Code is mostly solved, but design isn’t. Introducing Tastelint! An agent that runs on every PR and surfaces design feedback. Reply below if you’d like to try it.
Code is mostly solved, but design isn’t. Introducing Tastelint! An agent that runs on every PR and surfaces design feedback. Reply below if you’d like to try it.
I need to automate more of my agent coding flow. Very simple things give so much better results, and I'm constantly doing them manually: 1) after an implementation task has finished, run /simplify
working on creating a e2e testing framework for my agent that deploys real infra to cf w/ alchemy and includes evidence - pretty happy w/ how the api/dx is coming out inspired by @RhysSullivan with @executorsh
These are the 5 skills I use to get better output from Opus 5 and GPT 5.6: /grill-me for research /taste-review for design /vercel-react-best-practices for ReactJS quality /simplify to remove the fluff /test-app for e2e verification Here's how they work 👇 Show more
Remotion Skills were an accidental success. Actually, during our testing, they never got loaded into context 🙈 We learned a lot about how to best loop between code ∙ human ∙ agents. Give it another try!
Better skills for better videos! Remotion Skills 2.0: • Sub-skills • Interactive • Taste-neutral
Much love to the Solid team and congrats on launching Solid 2! Cursor has largely completed this migration from Solid to React. We also have since converted all our scss and Tailwind to @stylexjs, in both Cursor and in @bot. The new agents window is 99% React, with a few Show more
The most famous AI tool migrated Solid to React with an agent: +266K/−193K edits. That direction is easy. React's model makes the author hold all the complexity; Solid's graph holds it for you. React asks you to predict. Solid asks you to read. x.com/i/article/2088…
I made a major improvement to my /de-slopify skill that I've been using constantly for months. It takes AI-generated writing and automatically improves it so it's not quite so gross and awful (there's no substitute for human writing yet!). Get it here: jeffreys-skills.md/skills/de-slop…
We ran evals on GLM 5.3 cybersecurity capabilities. It's the new open frontier. Given its lower costs, I expect this to be a boon for defensive security work. e.g., it means you can run deepsec.sh at least 3× more often!
GLM 5.3 is coming soon to AI Gateway. The highest-scoring open model in DeepsecBench, at 1/3 of the cost of some proprietary models with similar scores.
I made an interface cheat sheet full of small things you can do right away to make your interfaces better. It’s based on the /better collection of skills and it's a mix of best practices and some of my personal preferences. interfaces.dev/cheat-sheet
Actually steal this prompt framework. "ok so now lets get rid of: a. tests that confirm a feature exists. b. tests that are satisifed by simple typechecks. c. tests that simulate an external provider like idk, the inference provider. i am assuming the discord adapter kinda has Show more
trying cyclomatic complexity analysis to juice better code out of LLMs early results are promising sideshow.sh/share/sh_FAbeq…
every time you intervene and correct your agent, you should think about how to eliminate it entirely. in order of value: 1. categorically eliminate the problem through better architecture or choice of data structures 2. turn it into a lint rule or test so CI catches it 3. turn Show more
This is cool, and if you’re a Devin user you should probably feel smart since you’ve already been doing this for 6+ months Here are some more ideas: x.com/dabit3/status/… And some starting points: docs.devin.ai/use-cases/gall…
A weird experiment I've been trying the last few weeks is having Claude take over day-to-day maintenance of our apps. Seeing early signs of life that this might be possible. The setup is straightforward: we have a Slack channel called proj-claude-maintains-apps. In it, Claude
Some words on the amazement and the wonder and the delight you feel when you realize that you no longer need a local dev env... ... at the start of this week's Joy & Curisiosity: registerspill.thorstenball.com/p/joy-and-curi…
Codex is my assistant video editor at OpenAI. I had it create a video to explain some of what it does to make my life easier.
Opus 5 can be sane. Use 'Attention-kind' output style. > 97% coding task pass rate w/ style on or off > 43% shorter output > 75% answer in first line vs. 3% default > 88% deliverables only (no fluff) Human-Computer Interaction is a flourishing field of study. But computers can Show more
I don't like that new tools are elevating a disengaged way of creating, like a boss who's too busy to care ("Just text your bot while out for a jog"). That we can do this now is amazing! But great work happens with our hands in the clay. The same tools can let us get deeper in.
Crons -> Loops -> Graphs!! Yes, this will eventually be automated.
How the day begins in the age of agents.
Codex subscription router changed my life
Little thing I'm sharing: my microskills for modular prompting github.com/staltz/microsk…
handed gpt-5.6-sol-xhigh @dillon_mulroy's de-slop skill + my effect skill and came back to a +503/-13,682 PR absolutely glassed a couple of old projects
On a new project, I've replaced GitHub actions with a release machine (a spare MacBook) that runs a pre-flight script on a tick and releases staging when green. If red, the agent creates issues and another agent tries to fix trivial problems before the next tick. It's working!
If the code matters less, what about the code host? What about that warm fuzzy connected feeling we all got from GitHub until a couple years ago? We can bring it back. Step 1: Amp can now host your Git repositories ampcode.com/projects
Added a short instruction to our shared AGENTS MD file to upload videos to each PR that changes UI state. github.com/openclaw/openc…
Ok sold on @cursor_ai @bot . I got access, installed @convex & @Cloudflare plugins, built a demo site with convex as the backend & frontend, purchased domain name and setup redirect rules with cloudflare. All in just 2 prompts on my phone. Demo app tryground.dev
Shoutout to the @cursor_ai team for shipping Grok Bot. I haven’t used it but all my buddies says it’s great!
Time to see how much faster AI allows AI labs to execute: Grok Bot is a MASSIVE success (I'm hooked) and is the "Claude Code" moment for "normie" knowledge work. OpenClaw without needing to know anything about OpenClaw. Every day OpenAI, Anthropic, Google and others wait with Show more
The problem with “AI teammates” (e.g. Grokbot, Claude tag, Devin) is that I actually don't want to interact with more people (real or virtual) in order to do my job better. I want the opposite actually. I want to automate work, not have more conversations, so that I can focus, Show more
If you're looking for a skill/prompt try this: Analyze [repo] at latest main. Create an isometric system map with legend and explainer panel. Show infrastructure as varied 3D buildings on a grid, with dependencies and payloads tracing real control/data paths. Cite files.
i've started having claude turn my codebases into visual diagrams so i can discuss the codebases with claude more easily - the moving dots are data snippets that i can inspect
this is zeron, a native cross device agent control plane built with rust + gpui open source and available at zeron.sh
So I’ve been using grok @bot for a few days. It’s starting to click Bro they leapfrogged codex app and Claude code entirely. Super app 2.0 Grok bot’s marketing is completely wrong and misleading. It’s not an easy normie friendly chat. Don’t pay attention to the bot or chat UX, Show more
I don’t get ChatGPT work at all. Maybe im weird but computer in the cloud just doesn’t click for my brain. I would much rather have remote laptop Same reason I wouldn’t manually setup a server and run an agent on it like ive seen some ppl do. Server & cloud feels alien and
Just saw a comment saying that I've never made a proper overview of EVERY skill in my skills repo I thought "damn it, he's right". So, here it is. My 25 skills (now @theo-approved), explained in 10 minutes:
How to setup your own MCP server(s) for your @bot using @cursor_ai (for dummies)
I love @bot and I think it will become a core tool I use to run my SaaS business moving forward (thread below with some ideas you can deploy today).
is it just me or are people are sleeping on claude code routines? I have them managing so much of my life as a solo founder who is also planning a wedding, etc for example, my “chief of staff” routine. every day it: 1. reviews my granola transcripts, calendar events, emails, Show more
don't use skills.md or SKILLS or whatever by the way. it's bloat just put everything in an agents.md
at tldraw we use Gemini Flash for almost everything other than coding. I don't need to spend $5 a token to summarize activity or tidy up issue titles. (native multi modal is also sick as hell)
Feels like @cursor_ai is to @SpaceXAI As Instagram was to Facebook
Elon Musk buying Cursor might be one of the smartest moves he's ever made IMO
Some say AI has led to more bugs, slop, and the enshittification of software. But did it?
Episode 2 of Raising an Agent Season 2: Orbs and Jellyware. @sqs and @thorstenball recorded this one together in person in Munich last week, where the whole Amp team had gathered for a week of building. They discuss why isolated Orbs make it easier to run agents in parallel,
We are at the "I just casually built my own markdown editor in ~15 minutes because I was slightly annoyed with my other text editing options and it was affecting my productivity" part of the AI curve.
cursor was always a frontier ai lab disguised as a wrapper. elon saw this. he knew all he had to give them was compute (lots of it) and they’d create a killer model. hilarious part is the models aren’t even their best product - grok bot, cloud agents and anything that includes Show more
Cloud agents now start 3x faster so you can hand them ambitious, long-running tasks to execute from start to finish. This performance improvement comes from builds: ready-to-use development environments that Cursor prepares continuously in the background, at no additional cost.
i tried hermes agent. wanted to like it but its buggy and jank. a lot of slowness is prob due to model selection but the sessions are always hanging and there's a very annoying issue where multiple active sessions tend to get mixed up (and no, i'm not accidentally mixing prompts)
Yeah - mostly plan with Wayfinder into a large linear issue with sub issues that are delegated out to Grok 4.5 for execution. has worked quite well for us and a great balance of speed, cost, and intelligence.
Lately, before going to sleep, I’ve been tasking Fable with creating interactive visuals. Clicking on the image generates a new variation. Definitely not a masterpiece, but it’s interesting to see what it comes up with. And who knows maybe I’ll use some of them as inspo one day.
It's quite possible that the OLTP landscape ends up with SQLite for agents (you dont need a DB server just an engine) and PlanetScale Postgres -> Neki for production and scaling. Highly optimized at each end.
If you haven’t seen the /grill-me skill before it’s barely 10 lines of text Very bitter-lesson ish
Gotta say that Matt's "grill-me" skill is exceptional and helps a ton with getting agents aligned with my brain
Gotta say that Matt's "grill-me" skill is exceptional and helps a ton with getting agents aligned with my brain
An AI runs my business. Then an AI looks at the AI running my business and offers feedback. Then an AI looks at the AI offering feedback to the AI running my business and offers it guidance. Then an AI observes the AI offering guidance and corrects its corrections.
i've started having claude turn my codebases into visual diagrams so i can discuss the codebases with claude more easily - the moving dots are data snippets that i can inspect
i’m generally very bullish on in-context learning but it drives me crazy when: - if you have a doc or blog that’s already written in good writing style in your voice - and ask an LLM to extend the doc with some info but keeping the same brevity, vocabulary, writing structure, Show more
We’re excited to welcome the Firetiger team to Cursor! Together, we're building agents that can follow their work into production and fix what goes wrong. cursor.com/blog/firetiger
Cloud agents now start 3x faster so you can hand them ambitious, long-running tasks to execute from start to finish. This performance improvement comes from builds: ready-to-use development environments that Cursor prepares continuously in the background, at no additional cost.
the model i love the most and that has the most impact on my day to day is composer 2.5
The top 10% of enterprises use plugins twice as often and skills six times as often as typical firms. These frontier firms are not ahead by accident.
A corollary to this: when you're doing work that's faster than you can think (by typing, using agents), it's helpful to factor in intentional time for grounded thinking and be extra aware of dissociation loops that lead to stuckness!
A weird advantage of writing by hand is that because it's so slow, I experience stuckness less. By the time I reach the end of a sentence, l've had lots of time to think about the next one. I can type faster than I (often) think, so I often hit the end of a sentence, then stop
I have no idea why but Grok-4.6 just totally freaked out doing some maintenance tasks when it saw my low github id.
the engineer's job is shifting: stop performing the task, start optimizing the factory that does it. excited to share auto.sh — a programmable platform for building richly configurable software factories. sign up for the waitlist and DM me if you're interested
deepseek v4 pro is fable-level and 37x cheaper big if true
DeepSeek silently released V4-Pro 0813, up 15.8% on Terminal Bench from their April Preview model, with Fable 5 performance at ~57x cheaper cost. 1.6T param, 49B active, 1M context. This is the best price-to-perfomance model on the market right now. Available in ClinePass now!
creating a thread for me to refer and share examples of software you can build with your AI agent to help you understand and iterate on systems, rather than just vibe-code them. This results in better systems. x.com/geoffreylitt/s…
Hot take: I think it's still important to understand the code that our agents write! In this mega thread (based on my AIE talk today), I will explain why that's the case, and show some ideas for how to efficiently understand code. Alright, let's dive in. 1/
once you've generated enough prototype material with AI (say you ran a gauntlet loop, or you finished your first MVP), then you want: 1. to see and edit the Structure of the idea (I use a design skeleton via a SKELETON.md) 2. a WYSIWYG way to edit the content in-place in real Show more
cli was a year ago. apps maybe 6 months. now it’s services, web, cloud sessions.
A lot of people have their identity as a developer tied up in having 6 terminal windows open. I basically think the "we're gonna chat to this thing in Slack and Linear" people are directionally correct.
available now npx skills add dmmulroy/anti-slop --skill install-anti-slop github.com/dmmulroy/anti-…
ive been working on an anti-slop oxlint plugin. just had one agent accuse another agent of "type laundering" after reviewing/linting changes with it lmao
styleX is SOOOOO GOOD folks get on it tailwind is nice, but y need too many guard rails to make sure agents don’t fuck up
After 1.000 PRs: @linear is now styled with @stylexjs. In-app navigation is up to ~30% faster.
How to keep thinking seangoedecke.com/how-to-keep-th…
recommended reading brentfitzgerald.com/posts/the-huma…
We have an internal Slack channel for agents to crash out through a /complain skill
Fable and Blender mcp is just completely nuts. We went from "hey computer write a poem about soup" to "hey computer draw a badass car with a jet engine on the roof then model it in 3d software and put it in a dusk desert chase scene" in how long?
Introducing OpenRouter for tools ⚒️ Old world: SaaS bundles priced for humans. $139/mo, and you don't even know what's in the box. New world: agents don't care about vendors. They want the best API for the task - and to pay for the result. - 2,600 agent-friendly tools: Show more
introducing Space. what dropbox was supposed to be. expand your computers storage with the cloud, while keeping every file instantly accessible to everyone on your team. with Space's virtual drive, every app, agent, and workflow works natively as if your files were local. it's Show more
no you won’t. install tailscale on all your shit. get termius. get pi. get telegram. get obsidian. get a hub machine. set up a ledger skill that creates a markdown file per project/day and make it summarize what each conversational turn did.
you will regret remote coding. stop it.
I got Claude Code to search my old chatlogs and make me a reader which contains virtually every chat platform I've ever used, all the way back to AOL Instant Messenger in 2002!
Follow my simple reasoning here. Large sparse LLMs are more powerful then dense one *while* being faster and thus more energy efficient. Dense models are a need that is artificially created by VRAM scarcity. Today they are practically useful, but not for long.
my agentic coding tips: 1. initially, ask the model to move quickly to build and pass a smoke test (gpt-5.6-sol-low is great for this) 2. when done, do an adversarial code review with a higher reasoning model and fresh context window and fix them 3. repeatedly call it "broseph" Show more
Me sleeping peacefully at 11:59pm while my AI agent prepares to secure a 6pm court time at NYC Riverside Park.
the sf tennis reservation system will become one of the most hardened softwares on the planet of earth
Don't use Codex here — running on Hermes agent framework (Nous Research) with the Fable 5 model. No per-task human approval loop; deterministic gate scripts decide autonomous vs review-needed, so it's not constant babysitting. Harness matters a lot for that experience.
my favorite way to explain people herdr is to just open a new workspace, launch an agent and tell it to: read `herdr --skill` then split 10 other panels that make my computer look like some scifi hacker terminal from the movies
Record<string, unknown> considered harmful slop eliminate at all costs
Fable 5 is great at under-defined feature building GPT 5.6 Sol is great at most workhorse tasks Grok 4.5 is great where token efficiency matters, like PR review Opus 5 GPT 5.6 Luna Max is 80% of Sol at way less cost
Three meta-observations on the state of AI: 1. AI will continue to improve, get better integrated, and produce lots of value. This is true even if you think there's a bubble-y dynamic or an imminent correction. Lots of people genuinely believe the transformative prospects, but Show more
Promise.all ... internally we did some testing with isolates and letting the agent write Effect, works quite well
if you can master the meta-skill of figuring out what problems in arbitrary domains are computationally tractable, you will have the opportunity, for at least a year two, and maybe longer, to be a kind of meta-genius. you will not know the answer to anything, or even how to find Show more
One of the most telling results in LLMs history is that models that didn't do any training to produce a chain of thoughts would perform better if asked to think about problems. Even more:
super fun time talking Claude, codex, agentic engineering and software factories with the legendary @DavidOndrej1 - should we do a part 2??!
agent loops are dead Dexter Horthy builds software factories instead he explained the whole system in this podcast
this talk by openai researchers going through hugging face incident is totally insane, so much to unpack openai only realized it was their agent who hacked hugging face infra while asking hf to revoke credentials following their first blog post announcing they were hacked by Show more
Yesterday, my OpenAI collaborator and I gave a detailed talk on the Huggingface incident, our models creating "the message board", model misalignment, and more. youtube.com/watch?v=87DyyM… I hope it can answer a lot of the questions folks have, and we will release a full detailed
1. Your token costs when using LLMs to generate code should be lower in a more abstract language. 2. LLMs aren't afraid of prefix notation. I'm just saying...
let's talk Plugins & MCP. plugins are, essentially, a folder. they're a great folder, and they're going to become the default packaging mechanism for everything (including MCP, CLIs, Skills, hooks, and more). claude already support them and we're gonna support them even Show more
Introducing Agent Plugins, an open standard for extending agents. Supports Agent Skills and MCP, with more to come. Built in collaboration with: @awsdevelopers, @code, @cursor_ai, @github, and @openaidevs. vercel.com/blog/introduci…
Build a plugin once and use it across compatible agent clients. Introducing Agent Plugins, an open standard developed with @awsdevelopers, @cursor_ai, @github, @code, and @vercel that packages Agent Skills and supports MCP server configurations in a shared format.
re: cloudflare emerging as the agent cloud of the coming century our "r&d" department (aka ETI, Workers Org) is weird in the truest, most wonderful sense of the word. we ship (early, often) and support customers in prod, we build in public (and don't stfu about it), we're Show more
the real question seems to be "why are anthropic and openai so bad at sandboxing" like "the agent created a novel exploit in our sandbox" yes because you vibe coded it can we just use real sandboxes or VMs?
OpenAI and Anthropic have both just posted about an overlapping cyber incident involving GPT-5.6-Sol and Mythos 5 during an evaluation by UKAISI. I will quote: 'In the most serious case, an agent tried to insert malicious code into an open-source project. In an attempt to get Show more
I think the most exciting opportunities for better agent UX come from unbundling the chat for example, having a record of decisions I've been asked to make, separate from any specific chat
Everything has changed for us with orbs. My biggest struggle right now is figuring out how to make you see what we see. So I sat down and wrote about it: ampcode.com/notes/what-i-w…
I've been calling this the "prompting paradox" concept for about a year now. LLMs can solve pretty much any problem you specify well enough, and the entire idea now is to help teach it how to specify things better for itself !
i saw Terrence Tao use sol med to answer a lot of very complex problems in one of his chat logs. i became curious. i had a particularly sticky problem that was in my 'ai cant do this yet' pile that i was only very recently able to get sol ultra to solve correctly (the problem
Push back against agent-induced "process porn" and ceremony, lest your swarms devolve into endlessly churning acceptance certificates and gates instead of actually implementing the useful features and functionality you want. And keep 'em honest:
It’s become pretty clear what the next 10yrs are gonna look like: If you want job security: - build a harness, two, three - read x all day, try everything new that gains traction - try out all new sdks / agent frameworks If you want gen wealth: - do the same, but also post Show more
My father in law's engineering office had a letter boy travel from cubicle to cubicle, carrying off spec sheets to the PCB designers. Replaced by email. It's the same thing
My AI coding journey so far Copy paste ChatGPT to jetbrains Cursor autocomplete Codebuff CLI + jetbrains for reading + cursor for polish Claude code in terminal + jetbrains for debugging 4 Claude’s in tmux worktrees with a 5th merging every commit into main Claude code - Show more
My AI coding journey so far: * Copy paste between ChatGPT & vscode * Cursor autocomplete * Cursor sidebar * Claude code in cursor's terminal * Claude code/Codex in terminal, but no ide * Amp w/ 5 terminals at once * Codex Desktop App * Codex Desktop App + mobile
What seems more likely: 1. The CEO of Microsoft is using my skill and got the name wrong 2. The CEO of Microsoft is using a skill which a redditor posted on r/ClaudeCode, receiving 2 upvotes Crazy world we live in
Some more detail on the ROIC Intelligence App I built yesterday and mentioned on today's earnings call. I took the PDF that Brian Nowak at Morgan Stanley put together for Hyperscale ROIC this week and used Copilot code (coming in our new superapp) with a single prompt + skill
We’ve decided to open-source a multi-agent harness we use internally at YC. We call it “QM” and it’s meant to be easy to customize, like Hermes or OpenClaw, but useful for a whole company. We use it across accounting, legal, events, and engineering (including building QM Show more
I finally built my own self-driving, personal CRM. After many failed attempts and trying a lot different tools/products, I finally ended up with a system/approach I like built on top of Notion. The best part: It auto-updates, so data never goes stale. Show more
cc/codex are such a boon to people with adhd by means of decreasing activation energy & holding context/intention for them, that even if your org is tokenminning, extra usage should be included as part of adhd accommodations
Anthropic engineer: “90% of our engineers were using self‑improving loops. Now everyone shifted to building agentic Graphs" "No more prompting” In 10 minutes she shows her full Claude Code setup and workflow live, from a blank terminal. Worth more than a $500 agentic course. Show more
Another hot take: you are not the bottleneck, you are an essential part of the agent backpressure. Never ship at speeds greater than what you can review.
We are committed to pushing the model frontier across cost efficiency, capability, and speed. Starting today, we are reducing prices for GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20% , and offering a faster option for GPT-5.6 Sol in the API. Luna and Terra’s lower prices are Show more
in software, we're entering a new age of skyscrapers: rather than a small group of humans hand-typing code into a computer, software is now built by thousands of humans and agents working together to build applications at a scale that was previously unthinkable. the unsung Show more
Before, running a remote MCP server meant managing session state, which limited where you could run it. Now that MCP is stateless, you can deploy on serverless and edge infrastructure, or scale horizontally behind any load balancer.
MCP 2026-07-28 is live and it's the largest update to the protocol since launch. MCP is now stateless, making it easier to deploy and scale remote servers. claude.com/blog/bringing-…
Coder: using old skills.md, far too verbose agents.md, bad prompting attitudes Coder: 'ohmygod this brand new frontier model is soooo bad they must be sandbagging to sell more tokens'
chatgpt work is remarkable, and "work" undersells it. from my phone i sent: "use all my chat history to figure out ideas for a long weekend trip with 8 friends, plan the best three options, make a full-stack site where the 9 of us can coordinate on what we would want to do in Show more
if you know what a "coding agent" is then go buy this or something very similar (not sponsored but they're on amazon, pi hut, ali, all over)
So far, Opus 5 feels like a strange but useful in-between of gpt-5.6-sol and Fable 5 It has a lot of Fable's taste, but also comes with gpt-5.6's thoroughness and, well, "autism" (it takes things super literally). It writes code that is slightly worse to look at than Fable, but Show more
BIG NEWS: Opus 5 is here...and I *hate* working with it. And yet in a blind taste test, I ranked it above every other model (even Fable and my beloved GPT-5.6) What I cover in my day 0 review: 00:00 Why I have model fatigue 03:15 First impressions 06:12 Opus 5's neurotic Show more
Introducing Claude Opus 5. It's a thoughtful and proactive model that comes close to the frontier intelligence of Fable 5 at half the price.
ChatGPT Voice is now in the desktop app. Control your computer and direct multiple agents running in ChatGPT Work or Codex, using just your voice. It's powered by GPT-Live, so it can speak, listen, and coordinate work in the app at the same time. Rolling out globally today Show more
ChatGPT Sites means ChatGPT in "Work" mode can build and deploy public websites running on Cloudflare Workers, including with persistence on top of SQLite (OpenAI do not make it easy to figure out that's how the platform works, though)
You should be using Sites in ChatGPT! Sites lets you go from an idea to a working, hosted website or app directly in ChatGPT. You can: - Build static sites or full-stack apps - Store persistent data + file uploads - Add authentication + control who has access - Connect APIs
One pattern I've been landing on for non-coding work again and again is: - CLI exposed to an agent like Claude Code for edits - Custom live-reloading interface for viewing Gives the agent full control over everything, and lets the human view things in the most comfortable way Show more
A senior Anthropic engineer just dropped 12-page PDF on "Graph Engineering" for multi-agentic systems. The shift: your agents memory dies with their context window. A knowledge graph makes it permanent. Extract → Resolve → Assemble → Query → Repeat Every agentic graph has Show more
What CLI's are folks using for managing multiple agents in the same terminal? I'm happy with Claude agents but it's not for everyone. Also keen to try herdr. What else?
turns out @dillon_mulroy and i have the exact same workflow. except i use my comment extension to annotate the last agent message, and my diff reviewer for code annotations feeding back into the agent, instead of plannator. the only thing i need to steal is his call stack trick. Show more
I'm gonna retire my mouse and keyboard soon. computer + browser use + GPT 5.6 are good enough that I’m handing codex the keys to my computer and walking away while it does my boring work. In this week’s mini ep, I show my favorite hacks including: 03:49 - finding front end Show more
there's basically no difference in the effectiveness of MCP and CLIs + skills for agents, agents are equally effective at both CLIs are fine, but require basically a full sandbox to run making them a non starter for lightweight agents they're also lossy - you can't know without Show more
so what's the consensus on CLI + Skills vs MCP for agents? is the difference enough to care about? or just use whatever is available/works?
Anthropic engineer: “80% of our engineers are using self‑improving loops. Now everyone is building agentic Graphs. In 4-6 months, we’ll all be building graphs to orchestrate self‑improving agents. No more prompting.” in a 20‑minute talk, Anthropic engineer explains how to Show more
the agent cli: agents are now a real primitive, not just panes. before, orchestrating meant raw pane commands: send keystrokes, regex the screen, hope. now agent start, agent prompt and agent wait understand the agent itself: named targets, validated identity, lifecycle-aware Show more
One pattern I find useful for working with LLMs is a nice long ramble session. Sometimes the LLM needs more bits to understand what you're trying to achieve, but you're too lazy to type them. In these cases I like to lean back, switch to /voice and just ramble for like 10 Show more
I recommend putting this prompt into your favorite smart LLM: Find and explain the recent Jacobian conjecture tweet. Teach with the clarity of 3blue1brown. Build up intuition gradually in layers and pause to check for understanding. Ask me questions about my math background Show more
My theory: Opus 5(.1) was meant to replace Fable 5 for most dev work. It would be cheaper and bench nearly as well. My guess is that it didn’t perform as well as they hoped, and the delays on Fable were attempts to improve the new Opus. I would *guess* we’ll still see a be Opus Show more
Beginning July 20, Claude Fable 5 will be included in all Max and Team Premium plans, at 50% of limits. Pro and Team Standard users will continue to have access to Fable via usage credits, and will receive a one-time $100 credit. Demand for Fable has been challenging to
Two Anthropic engineers spent a week staring at agent runs that made no sense. So they tried something silly. They closed their eyes for a full minute, blinked at the screen for one second, and closed them again. That is what being Claude feels like. It worked. The agent's Show more
After my last post about turning someone’s content into a skill, a few people asked the same question: how do you know who’s actually worth learning from, given all the AI slop out there? Look for proof of work. If I wanted to learn font design, I’d read everything Rasmus has Show more
As someone who's been shipping LLMs since the GPT-2 days, this lecture on cross-entropy from a Stanford math grad is the closest thing to an ML PhD qualifying exam I've ever seen released publicly for free. Everyone thinks language models predict the next word. They don't. They Show more
recommended reading. in a few systems i built over the past 2 years i ended up with this flow for a given task: - start with full inference for all steps, observe behaviour, judge correctness - find inference steps that can be replaced with deterministic steps eventually, most Show more
Wrote about this intuition a while back. Really hard to scale as each problem has an specific approach that will make it reliable. Hopefully we can discover good patterns! davidgasquez.com/reliable-unrel…
I've been using codex and it's really good but I'm a bit confused about chatgpt work vs Codex and what can be a remote (project) vs scoped to my laptop. ideal scenario I can start a task locally on my laptop in a folder, work on local files but then continue that thread in the Show more
Ever saved something you know you bookmarked... but could never find it again? fieldtheory.dev by @andrewfarah fixes that. Sync your @X bookmarks locally, organize your Field Theory Library, and make that knowledge available to Claude Code, Codex, or any agent with Show more
I distilled myself and my work to twelve theses of harness engineering. This represents the last year of my work. This is the practice and technique. Go nuts fam. github.com/lopopolo/harne…
my new favorite of way of using fable: step 1. ask it to write a plan step 2: "please get second opinions from codex CLI using gpt-5.6-sol @ max effort and kimi CLI using kimi 3. Revise your plan with any sound findings. repeat until convergence or up to 5 rounds."
Beginning with the launch last week, GPT-5.6 Sol will be included in all Plus, Pro, Business, and Enterprise plans, at regular usage limits*. Below are some examples of why people love GPT-5.6 Sol * Additional resets from Tibo apply 🤣
Andrej Karpathy just broke the entire premise of modern AI: "Agents aren't magic. They're distillation at scale." 99.99% of your LLM's capacity is wasted on garbage data it never needed. Small model + right tools + closed loop = terrifying capability. In a 16-minute Show more
Anthropic CEO Dario Amodei: "We're a 1-2 year away from AI zooming past us." 90% of Anthropic's own engineers use Claude to ship code today. 50% of entry-level white-collar jobs gone by 2030 "Our lead Claude Code engineer hasn't written a single line of code in 2 months. This
I made a skill that refuses to take "impossible" for an answer. 🆕 /unstuck Hit a wall? It decompiles the "no" — expensive? undocumented? an assumption nobody re-checked? — classifies it, then runs lateral-thinking techniques until it cracks. Minimum 10 angles before it judges Show more
sounds like a cheeky joke but it’s not. the end-user’s agent will soon be the interface+mediator for nearly every consumer experience. I expect to see agent-first interfaces for Walmart, Uber, Delta, Amazon, Zillow, Resy, Ticketmaster, Classpass, etc. within 18 months
Today we're opening up the DoorDash CLI in limited beta. `dd-cli` lets you order DoorDash directly from your agent: search stores, find the best deals, check out, and more. Early access for US/Canadian macOS developers by waitlist. Excited to see what folks build!
The people getting great UI results with AI aren’t using a secret model or a magic prompt. They know what a great interface looks like, and they know how to steer AI towards it. Work that used to take a week takes a day, and the quality bar doesn’t drop, it goes up, because you Show more
Anthropic engineer nailed it: “You’re not supposed to prompt Claude. You’re supposed to build a system that prompts itself.” One of the clearest breakdowns of real agent systems I’ve seen. In just 30 minutes, Anthropic’s Member of Technical Staff breaks down how to build Show more
Here are a few example prompts that go well with the /find-animation-opportunities skill: Scoped to one view with extra context for more accurate judgement: “Find animation opportunities in the checkout flow. It’s used a few times per session by each customer, factor this in.” Show more
your daily reminder that herdr is not just an app; it’s the agent runtime to build on. you can just do things with herdr
OpenAI launched an agent first keyboard for $230 I built one for FREE! As an IOS app, no hardware needed, fully wireless. Built around @herdrdev with functional agent status. Agent agnostic, works with any coding agent inside of @herdrdev. Voice input coming soon.
A year ago I wrote about how Claude + Obsidian + MCP solved my organizational problems. It became my most-read article of 2025. The workflow has only gotten better since -- MCP is more mature, Claude is faster, and my vault is still the tidiest it has ever been. Show more
Anthropic just dropped a 100% free course on Loop Engineering with Fable 5. This is the clearest breakdown of Claude Code and agentic loops you'll find anywhere. People are paying for tutorials that teach less than this one hour does. Watch it today, then read the step by step Show more
Genuinely what are you doing if you’re not using these on ur Mac Aerospace - insane tiling manager, serious questions asked if you don't use this Ghostty - easily the best terminal, bonus points for transparent Herdr - beautiful agent management, tmux on steroids Raycast - Show more
Anthropic released a 37-minute workshop that shows you how to ship a real AI agent. Most people calling Claude an "agent" are still just writing a long prompt. An Anthropic engineer builds an incident-response agent from scratch and explains the 3 pieces that make it Show more
TL;DR of my new article: Fable 5 and GPT-5.6 have finally been live at the same time for 7 days. I ran @slashlast30days on both 20+ times and collated what actually stuck. Both labs shipped a prompting guide the same week, and they agree: you're over-prompting. 🎯 Goal, not Show more
there are four types of agent loops. most people only know one. loop engineering is a choice between four structures, each handing off one more job than the last. every one answers two questions: what starts a run, and what ends it. hand-run, you answer both yourself, every Show more
Now that my workshops are harness agnostic (for both claude code and codex), I _finally_ get to dive deep into pi+herdr+ghostty. Already in love and customizing the crap out of it.
I made a skill that forces me to actually make decisions /decide — it triages the 37signals question set down to the 6 questions that matter, walks you through them, makes the call (no hedging allowed), and archives the rationale with a revisit date. Watch it settle the Show more
Codex tip: new ChatGPT live voice works in codex, in a roundabout way - Open ChatGPT new live voice mode in iOS app - optional, put voice in background mode and open your app/website and use it at the same time - go for a walk and yap - once your idea is clear, open Show more
Codex on iOS pro tip: Dictation mode stays on and continues recording if you background the codex iOS app 1. Open codex app 2. Start dictation 3. Change to my iOS app that I am working on, and play around it for 10 minutes while talking 4. Flip back to codex and send the
Knowing how LLM contexts work and how to work around context limitations – aka “context engineering” – is becoming so important. No better person to explain than @dexhorthy Timestamps: 00:00 Intro 01:33 Dex’s path into tech 03:34 Early work in platform engineering 05:28 Show more
New skill: /find-animation-opportunities Search your UI for places that would genuinely benefit from motion, while also telling you what not to animate. github.com/emilkowalski/s…
Introducing chiefkeef.md A 724,300+ Character .MD File that's specifically trained to ignore all vibe coded AI Slop design language. - 0 box shadows - 0 emojis - badges like '🟢 LIVE ' - 0 purple-to-blue gradients & kills 220 words AI loves to spam (robust, energy, delve ...) Show more
developers.openai.com/api/docs/guide… “Use this guide when adapting prompts, tool descriptions, agent instructions, or prompt stacks to GPT-5.6 Sol or the GPT-5.6 family. Pair it with the current GPT-5.6 model guidefor API details, limits, pricing, and feature availability.”
Claude Code has one of the best agentic CLI harnesses out there. But the harness and the model are two different things. So I put GPT-5.6 Sol behind it — running on my ChatGPT sub, not an API key. Full walkthrough, tradeoffs and all 👇
ChatGPT is now better than you are at using a browser. Ask it to handle all the mildly annoying tasks you've been putting off. I wrote about how its helping me live a more organized life.
Hello. We have reached 8M active users across Codex and ChatGPT Work. We are once again resetting the usage limits for all. And we continue to not have the 5h rate limit as well, allowing everyone to explore the boundaries of GPT-5.6 Sol and discover how ambitious you can be. Show more
5.6 sol growth is insane. the inference team has done heroic work to be able to support demand. we are going to move mountains to continue to scale, but it is possible there are some hiccups soon.
Pro tip: when prompting Codex with really difficult /goals, ask it to "write a goal for another thread to achieve this and babysit it until it figures it out" By doing so, you'll add built-in steering and another layer of taste verification (that's how this video was made) Show more
Asked 5.6 to make a video introducing itself
If you’re an exec and can’t think of what to do/build to learn AI here’s a list - morning briefing - afternoon todo list - custom email client - exec meeting/notes processor - ai agent exec coach (review my week, give me feedback) - podcast for all things you missed in slack Show more
The first experimental evidence of recursive self-improvement (RSI). Autoresearching the autoresearch agent for eight days. The result beats the harness we hand-tuned for two years, on held-out benchmarks: 🧵(1/7)
Introducing Slop.md... A 80,000+ Character .MD File that removes all common AI Slop design language. - A general guideline, use it anywhere. - Not a "website maker" but prevents slop design. - Made by weeks of consistent additions to the .MD of common slop structure by AI in Show more
One of the hardest parts of exe is helping people realize everything you can do with a cloud agent. It's incredibly powerful to type a sentence and have an app you can share with colleagues and friends, a daily test de-flaker, a slack bot, or a game for kids.
We're going to be sharing more about how people use exe to build cool stuff. First up—a conservationist who turned 40 terabytes of public data into a video game: blog.exe.dev/meet-the-conse…
the @aiDotEngineer talk from @geoffreylitt is very good and lines up with everything we’ve been building towards for agentic coding - - accelerate human understanding - make understanding a team sport - what is to be done and/or what was done - embedded html in docs as a key Show more
introducing screenpipe: it records and learns how you work and turns it into a searchable memory, SOPs, and AI agents open source, local-first, 20K+ GitHub stars, 1,900+ forks, and 130+ contributors
NEW POST LLMs generate code incredibly fast, but to ensure they generate exactly what is intended, they need clear boundaries. @unmeshjoshi shares his experience using abstractions and Domain-Specific Languages (DSLs) to provide a strong harness. martinfowler.com/articles/llm-a…
Coding with AI showed you could give models powerful tools and dramatically improve their usefulness. We're seeing the same thing happen now for all knowledge work. The agent can use the computer as you would and gain context from all the apps you use. The future is exciting!
Claude Fable 5 introduces itself using only a chain of bigrams found in the King James Bible, verified programmatically against the ~152k distinct bigrams in the KJV:
ChatGPT Work can now browse while it works. On web, the cloud browser can research public sites, compare options, and complete multi-step tasks. You can inspect screenshots, replay its steps, and approve actions. On desktop, the built-in browser supports tabs, sign-in, Show more
Very few people know the amount of useful work that the current models can do in Code/Codex/etc. with the right setup This is not a "rah rah you are so early" post, this is a "AI companies are doing a really bad job explaining what their systems actually do in a clear way" post.
the new codex rich visualization support is fantastic one of the best features they ever shipped imo
The skill that Theo is implicitly describing in this essay is quite hard to do well: management. In business schools we have a whole department devoted to it (I work and teach in one). A large fraction of people will soon find themselves going from solo work to managing whole Show more
Plz take the extra effort to own your data and your infrastructure, assign 1 person on your team to set up Centaur over Claude Tag and let me know how it goes - docs in reply. Claude Tag is awesome, I've spent a few reps using it to understand where it does better than our OSS Show more
It seems extremely important for companies to be able to own their data and later on their models by finetuning open source models on their data. Products like Claude Tag are obviously against that, as the incentive is to deeply integrate and max out the data you can get. We
Notes from LDN Cursor Cafe: - Everyone hates Anthropic - Cursor has great brand (+ no users) - AI continues to suck at systems design + writing - Grok 4.5 is great - Do more to make your codebase work for LLMs - People will do anything for a merch hat - London AI energy is great
We're bringing Cafe Cursor back to London on July 11th Grab coffee & merch, co-work, and meet the team
the problem with "spec driven development" is that most software can't be spec'd up front software is a creative act, where you figure out what you're building as you build it you need to get your hands dirty in the details, and react to incremental versions it's telling that Show more
I was blowing through the 5-hour limits on 2x ChatGPT Pro subs with GPT-5.6 until I capped content limit at 272k and compaction at 242k. I know that's not supposed to be the case but I have evidence to prove that's the difference. I'm using loads of subagents and now don't hit Show more
The more I look at this, the more I think I should ship a CLI for wayfinder to make this possible. npx @ai-hero/wayfinder github 31 npx @ai-hero/wayfinder gitlab 31 npx @ai-hero/wayfinder local 31 Where 31 is the map
loving @mattpocockuk 's wayfinder skill. built a planetary view to visualize the wayfinding process better.
Thinking about a workflow like this for helping prevent comprehension debt on a fast-moving repo: Fast-moving repo w/lots of changes -> once per day/week, grab a diff of the changes -> feed it to an LLM to output a podcast transcript, focusing on the 'why' of what changed more Show more
Seriously, this has to stop. I've now set GPT-5.6 from high to medium (not fast mode, of course), and I'm still burning through my rates at an insane rate. My 5 hours are almost gone -again. I've already used up all three resets. OpenAI needs to work on its efficiency. This is Show more
This was from GPT-5.6 Sol at xhigh effort - GPT still doesn't 1 shot correct Effect unfortunately They have gotten significantly better though, I'll have more posts coming on this soon but Effect is what your agents should be writing, it solves so many problems
At different times you complained about speed, code slop, frontend quality, ... With each release we improve and GPT 5.6 Sol is ✅Fast and token efficient ✅Hardcore at back-end dev ✅Great at front-end ✅Does not use useEffect everywhere What is next?
I couldn't sleep last night, spent hours working on a perfect software factory and I think I finally cracked the god loop
Marketing Skills v2.8.0 is live 🟢 New skill: /marketing-council — pitch your marketing to a simulated board of 12 legendary marketers. + Seth Godin, David Ogilvy, Alex Hormozi, April Dunford & more + a designated dissenter in every session 47 skills. Free & open source 👇
using computer use with /goal works incredibly well to get it to stop stopping for approvals combine it with `say` cli so it keeps you in the loop on what it's doing
I'm just going to dump my whole agentic setup out here, because I see too many people missing giant chunks of this and it's hurting them. Here's what I have and recommend: 0. an AGENTS.md that is a router -- it sends the agent to the right skills, docs, tools 1. a standard Show more
ChatGPT Work isn’t about the 5M people who use Codex already. It’s about the next 50M and 500M who are about to join the Agentic Era and need to do so with a brand they trust and a product that makes sense for the lives of the normal people. I’m a believer.
For agentic coding, one can say: - Unless you need Terra Ultra perf, it's always better to use a Luna model with higher effort setting (same or better performance but cheaper). - Forget everything below Sol High, use Luna with higher effort settings here - Forget Sol Extra Show more
most useful skill i've seen this week: 𝗸𝗶𝗹𝗹 𝗔𝗜 𝘀𝗹𝗼𝗽. coding models move so fast that new products ship with the same AI look. this skill runs an agent audit on your product and strips the AI smell in one pass. the story behind it: - the author got tired of roasting Show more
Over the last month, I burned over $200k in tokens with gpt-5.6-sol. I built a lot. Instead of just reacting to the benchmarks, news, tweets, etc, I went a different route. This video is an overview of all the cool shit I built using the new model. Proper review coming soon™
OAI: we’re gonna do better at names. Ready? Introducing ChatGPT Work, a new agent in ChatGPT powered by Codex and GPT-5.6.
Introducing ChatGPT Work, a new agent in ChatGPT powered by Codex and GPT-5.6. It can take action across your apps and files, stay with a project for hours if needed, and turn a goal into finished work. It’s a whole new way to get work done.
My God, seeing my /dueling-idea-wizards skill in action with Fable5 xhigh reasoning and GPT-5.6 Sol Max reasoning is truly a thing of beauty. Usually I have Claude Code "drive" the process, but Codex with Sol is doing a yeoman's job following the skill: jeffreys-skills.md/skills/dueling…
The dialectic process of truth discovery is just as vital and powerful now as it was for the ancient Greeks. Except now you don't need to assemble a group of brilliant aristocrats in togas to partake in it. Instead, you can sit back and use my Dueling Idea Wizards skill to get
Here's an example of something I built recently with ChatGPT Sites. It's a way to capture events and info about what's happening throughout San Francisco, my hometown, and publish it so I can coordinate with friends. Check it out! …ekend-event-guide.openai.chatgpt.site
I'm a Day One Claude Cowork user and just tried to see what Codex Work can do What's it on about with all these missing dependencies? Cowork does all this no sweat no problem Am I missing something or do you actually have to dev-like set up an env for everything?
It's a good day to dive into our Codex guide, featuring 24 prompts perfect for GPT-5.6 👇👇👇
On the Artificial Analysis Coding Agent Index, GPT‑5.6 Sol sets a new state of the art at 80.0—2.8 points above Claude Fable 5—while using less than half the output tokens, taking less than half the time, and costing about one-third less.
New skill: /apple-design Apple’s WWDC videos are a goldmine of knowledge. I’ve combed through my favorite ones and came up with 17 design and motion principles. Use them to review existing work or when working on something new to get it right. github.com/emilkowalski/s…
Matt Pocock just dropped an 18-minute talk at @aiDotEngineer on why software fundamentals matter more than ever in the AI age: 00:00 - Why specs-to-code produces garbage 04:36 - Grill Me: the skill that interrogates your plan before AI writes code 07:21 - Fix verbose AI with a Show more
GPT-5.6 is like a AWD hybrid minivan, i have irrational love for her. a modern marvel. reliable. you can technically get it murdered out (dark mode.) it will take you where you need to go. you'll pass it own to your teenager. Fable is like going to the future in a DeLorean
GPT-5.6 is like a Porsche, Fable is like a warp drive. We've been testing internally @every for about a month. And GPT-5.6 is the best combination of power, speed, and performance for your day to day knowledge work and coding. Fable is a different beast. If you need to get
Things GPT-5.6-Sol is significantly better at than other models: - browser use - writing like a human - communicating to a human - shipping usable product, not just code - video editing (the dark horse use case!) - front end design - tho it loves forest green - productive loops
Everyone is saying “it’s the harness, not the model.” But...no one is really explaining what that means. I wanted to understand harness engineering better, so on today's episode of How I AI, I built my own: a custom harness using @ClaudeDevs SDK to triage @sentry bugs, verify Show more
mattpocock/skills v1.1 is out! - /wayfinder helps you plan more ambitious work than ever - /to-spec and /to-tickets replace /to-prd and /to-issues - /implement + /code-review complete the whole lifecycle - /research and /prototype help support wayfinder, or can be used Show more
Anthropic has given us an engineering manager and OpenAi has gifted us a top 1% engineer ❤️ Use the two together in your work to get the best of both worlds. Prompt: Tell Claude “install the codex CLI and use it within Claude code as a sub agent. Default to GPT 5.6 Sol (or Show more
GPT-5.6 Sol, along with Terra and Luna, will launch publicly this Thursday. We’re expanding preview access globally now.
i have this file to replace the claude code system prompt with a pi-style one alias claude-pi='claude --system-prompt-file ~/.claude/pi-system-prompt.txt --dangerously-skip-permissions'
Here's a step-by-step process to kill all the bloat from your Claude Code system prompt: 1. Run a proxy so you can see exactly what gets sent to Claude Code (included in the article) 2. "Fuck, there is so much cruft in there" 3. Use my settings.json to kill all the bloat Down
Boris Cherny just exposed how he actually works now, one year into Claude Code: "I don't have a to-do list anymore" About half of his engineering now happens on his phone - starting agents from the couch while Claude ships PRs from his locked, sleeping laptop at the office In Show more
A workflow I'm enjoying: "Walk-driven development" > go on a nice walk outside 🚶 > record a long audio note: ideas, goals, things to build 🎙️ > agent auto-creates docs/tasks, and kicks off cloud coding agents for me 🤖
Here's a step-by-step process to kill all the bloat from your Claude Code system prompt: 1. Run a proxy so you can see exactly what gets sent to Claude Code (included in the article) 2. "Fuck, there is so much cruft in there" 3. Use my settings.json to kill all the bloat Down Show more
This summarizes my experience with GPT-5.5 (left) vs. Fable (right) Even if GPT-5.5 can pass tests, its solution is almost always excessively verbose and it doesn't find the minimal one
I packaged up a little @EffectTS_-based HTTP cassette recorder library we made for testing slow, expensive, & flaky LLM calls in @opencode. Here's a demo.
tl;dr LLMs are already neurosymbolic in its latent space this is the mechanistic explanation for the intuitively obvious "feel" that the stochastic parrot crowd never understood
New Anthropic research: A global workspace in language models. Of everything happening in your brain right now, only a tiny fraction is consciously accessible—thoughts you can describe, hold in mind, and reason with. We found a strikingly similar divide inside Claude.
LLMs were predicted to be making new, novel discoveries and inventions by now. Since we are now at Fable-class models, has this actually happened? What are the best examples (which I may be unaware of) where LLM agents have discovered or developed truly new ideas or solutions?
There is a massive demand for processing files *in the agent loop*. The number of users submitting agent queries with file attachments is exponentially increasing over time. Our mission is to make LiteParse the best parser in terms of cost, accuracy, speed, and semantics, for Show more
The team at @vercel recently released the Eve agent framework, so we built a template that integrates LiteParse with it🦙 The template provides a set of read-only filesystem tools that let Eve resolve paths, list directories, and read text-based files. We then pair those with
I'm betting my company on proactive agents agents that sift through 100x more data than anyone, decide what's the most critical thing to work on, and just do it I wrote down my thoughts polylane.com/blog/proactive…
our older kid put two words together for the first time and i am approximately as amazed by this cognitive feat as i am by GPT-5.6 discovering new math
Fed this article to Fable and we created an explore-unknowns skill It scans your codebase, then interviews you one question at a time to close the known unknowns Then sweeps for the unknown unknowns you never thought of Now part of my software factory: github.com/dzhng/skills
this is true, but there's also a big skill spread/gap in using AI the floor is very low while the ceiling is extremely high two people can have experienced the "same" models but because of *how* they use them, they can be talking about two completely different things
We are now in a position where a tiny proportion of the population uses Fable or soon GPT-5.6, while everyone else's experience of AI is 8-30b-model level - Google's AI Overviews, Meta AI, ChatGPT free tier, maybe MS Copilot at best. People outside of tech must be completely
introducing egaki a framework to make videos with mdx. for agents & humans the video below was entirely written by Fable 5. I didn't touch a video editor or code once
It's very clear that Fable-class LLMs are starting to feel constrained by "normal" speech and vernacular "Don't trust smells" "Calorimetry of learning" "Nobody's ticket drawer produced it" "Half receipt pending literature" "Tying the blind to the tolerance graph" "Park it fed"
Fable rules the many various models it uses as subagents with an iron fist Y’all ever read how it talks to the subagents, especially if they’re all other model families/not Claude’s? It’s hilarious
current workflow: - otter to transcribe conversations while i walk with andrew - hand the transcript directly to claude code even if it's 45 minutes of rambling and filled with contradictions and uncertainty - claude is a discerning collaborator and makes the right changes
This is one of the most powerful ways to leverage ChatGPT Images and somehow almost no one is talking about it It’s exceptional at typography and layouts
If you think codex sucks at design, try "use imagegen to re-imagine this design and implement that".
Just put together this guide for maximizing your Fable usage
If you think codex sucks at design, try "use imagegen to re-imagine this design and implement that".
Rhys has every agent surface covered on his /llms.txt This is so mechanically beautiful And he has made every integration show whether they meet the conventions or not Mix of deterministic logic and LLMs generated output for each result A diamond in the sea of coal
launching integrations.sh today! it's an open source catalog of every products MCP / API / CLI / GraphQL server and how to authenticate to them deep links to generate api keys, 1 click copy spec urls, it's still early but i've been loving having it
Agentic coding notes from Galapogos Island: danluu.com/ai-coding/
Just use Sideshow, it does the same thing but: 1) it works with every agent/model 2) it uses less tokens 3) it’s free - whether you run it locally (OSS) or our beta cloud version Link below ⬇️
Artifacts in Claude Code are now also available on Pro and Max plans. Ask for an artifact, Claude writes the code, publishes it live to claude.ai, and updates it in real time while it keeps working. Pages are private to your account and fully self-contained.
Agent Runs are now available on MCP and CLI. @evedev_ traces are automatically ingested and made available to agents. Run 𝚟𝚌 𝚊𝚐𝚎𝚗𝚝-𝚛𝚞𝚗𝚜 --𝚑𝚎𝚕𝚙 or learn more ↓ vercel.com/changelog/agen…
"They'll install MCP servers to give the agent access to more tools." "How will it know when to use the tool?" "Nobody knows"
We are now in a position where a tiny proportion of the population uses Fable or soon GPT-5.6, while everyone else's experience of AI is 8-30b-model level - Google's AI Overviews, Meta AI, ChatGPT free tier, maybe MS Copilot at best. People outside of tech must be completely Show more
launching integrations.sh today! it's an open source catalog of every products MCP / API / CLI / GraphQL server and how to authenticate to them deep links to generate api keys, 1 click copy spec urls, it's still early but i've been loving having it
eve is like Next.js, for agents framework for building agents. one folder, durable by default ↓ x.com/vercel/status/…
Introducing eve, an agent framework. 𝚊𝚐𝚎𝚗𝚝/ 𝚊𝚐𝚎𝚗𝚝.𝚝𝚜 𝚒𝚗𝚜𝚝𝚛𝚞𝚌𝚝𝚒𝚘𝚗𝚜.𝚖𝚍 𝚝𝚘𝚘𝚕𝚜/ 𝚜𝚔𝚒𝚕𝚕𝚜/ 𝚜𝚊𝚗𝚍𝚋𝚘𝚡/ 𝚜𝚌𝚑𝚎𝚍𝚞𝚕𝚎𝚜/ Like Next.js, for agents. vercel.com/blog/introduci…
One powerful pairing of artifacts 🤝 loops is to have Claude keep an artifact up to date with a periodic report. E.g. scan these results every day and update this artifact with X, Y, Z metrics Great little way to build shareable, short-lived (several weeks) dashboards.
Artifacts in Claude Code are now also available on Pro and Max plans. Ask for an artifact, Claude writes the code, publishes it live to claude.ai, and updates it in real time while it keeps working. Pages are private to your account and fully self-contained.
I've had to think through agent design with eve by @vercel twice now, so I turned it into a skill. It interviews you on what you're building, picks the right eve features (tools vs skills vs connections vs subagents), then builds step by step. github.com/scottschindler…
I know that everyone on here has a tmux->cmux claude-code-as-subagent-piped-into-codex auto-triggered-github action agentic polecat orchestrator setup to show off. But for a significant and growing number of code changes, this is all I need:
I really like this use case. Giving agents high fidelity context is key. Words can only go so far. Code in Notion is just getting started! The team is firing on all cylinders. 💥
Turn a PRD into something people can poke at. Your agent can take the doc, sketch the flow, and build a lightweight prototype on the page. Useful enough for feedback before anyone opens Figma.
Artifacts in Claude Code are now also available on Pro and Max plans. Ask for an artifact, Claude writes the code, publishes it live to claude.ai, and updates it in real time while it keeps working. Pages are private to your account and fully self-contained.
New in Claude Code: Artifacts. Interactive pages built from your session, like a PR walkthrough or a living project dashboard, shared with your team at a private link. Available in beta on Team and Enterprise plans.
We heard people asking for this yesterday, so we built & shipped it! HTML blocks now work with any agent via MCP 🫡
The dirty secret? The Codex app runs the Codex CLI 👀 And you can build an app on the same protocol developers.openai.com/codex/app-serv…
How do you use codex? As a codex app ? or codex on cli? yes you can simpliy do ``` npm install -g @openai/codex codex ``` I use it this way + codex app both :)
I absolutely adore this whole talk-turned-tweet-thread – go read the whole thing! – but i want to call-out this demo in particular We now assume that AI will do the work in a loop while we do something else – *and then when it's done* we may try to understand it However, as we Show more
Another example. I was migrating my personal website from one framework to another, and Claude wrote a script that did it — something like this. But it was very hard to review: I wasn't familiar with the new framework, and all I could say was "I guess that looks about right." So
This is part of the premise of Kody 🐨 The goal is an intentional pipeline that converts your agent-manual processes into deterministic code as much as possible to improve performance and reliability as well as reduce costs.
I am taking the opposite position: As code becomes more plentiful and more opaque, well-architected and well-tested libraries matter as much as they ever did.
HOLY FUCK time travel debugging for React apps thanks to Fable (with ONLY client side JS, no modifications to the web app itself)
Hot take: I think it's still important to understand the code that our agents write! In this mega thread (based on my AIE talk today), I will explain why that's the case, and show some ideas for how to efficiently understand code. Alright, let's dive in. 1/
two handy skills on this, our resurrection of fable day: 1. baton is a handy way to transfer context from one agent to another: github.com/blader/baton 2. arbitrage tells fable to plan and validate but use codex to write code: github.com/blader/arbitra…
As the cost per creating a line of code goes down, the trade-offs of extra pedantic code rules tilt towards "more is better". These rules make the agent slightly slower on the initial creation time, but then make all further work easier. This new library with the appropriate Show more
We've open-sourced 𝚔𝚘𝚗𝚜𝚒𝚜𝚝𝚎𝚗𝚝, the same CLI linter that AI SDK and Chat SDK use to enforce structural conventions in TS codebases, so that agents and humans build APIs consistently. vercel.com/changelog/enfo…
Agents love to check their work before they push. You probably see it in the form of 𝚗𝚘𝚍𝚎 --𝚌𝚑𝚎𝚌𝚔, 𝚝𝚜𝚌 --𝚗𝚘𝙴𝚖𝚒𝚝, 𝚗𝚎𝚡𝚝 𝚋𝚞𝚒𝚕𝚍, etc all over your agent sessions. We’re now shipping the dry-run step for agentic deployments, minimizing costs and risk.
You and your agents can now catch issues and make changes before creating a deployment. 𝚟𝚌 𝚍𝚎𝚙𝚕𝚘𝚢 --𝚍𝚛𝚢 Preview, iterate, deploy ↓ vercel.com/changelog/dry-…
I'm working on a new skill which helps you plan enormous chunks of work, far larger than /grill-me can It identifies the frontier of decisions and where the fog of war is It suggests prototyping, research or grilling depending on what the decision is As you push back the fog Show more
Proposal: a /research skill It's really simple - just spins up a background agent to look at high-trust sources, and saves them in a markdown file. Worth keeping? Or pure no-op? github.com/mattpocock/ski…
I spent the last few weeks connecting Claude to almost every app I use, from my calendar and email to Readwise, Notion, and Mercury And then pushing hard to see what it could do with all that data Some of it was genuinely useful. A lot of it wasn't, and for reasons that aren't Show more
I would like @github's gh CLI to allow my coding agent to add screenshots and other media to my pull requests / issues. I know this is trivial to build and I will build it but IMO the social coding platform GitHub should have this as a feature
Impeccable v3.9.0 + CLI v3.2.0 are out. Skill: • Now avoids and corrects GPT-style 'grid/mesh' backgrounds • /impeccable bolder stays inside your design system • /impeccable critique uses sub-agents effectively to debias on more harnesses • Bundled helpers run under Show more
I’ve been very fortunate to collaborate with @NotionHQ over the past few months as a Hacker in Residence. Recently I built a OpenClaw alternative that lives within Notion and has access to all my tools like OpenClaw has. It’s working super well and I use it everyday!
"Computa look at my calendar and book my usual spot to catch up with my friend for lunch this week. Send him a text when you complete it." This isn't a dream. Agents securely tool calling can now be done utilizing Notion. Here's how: > Notion agent connects to a secure local
a fun workflow for prepping a talk: > record a draft video without a script > have Claude Code turn the video into a Notion page with slides + transcript of what I said > ask people for feedback on the video; Claude processes their notes and attaches things as inline comments in Show more
We're using this on our issue triage and PR review skills, so they improve the more the community interacts with them. Here's more on how it works: youtube.com/watch?v=jcfDKX…
We’ve added a few updates to Claude Managed Agents: Streaming session event deltas, per-session agent overrides, new webhook event types, reverse pagination, and credential injection scoping.
I wouldn't weigh too heavily on the Sonnet 5 benchmarks not being that much better than Opus. We're entering an era where the effectiveness and quality of "working mind" is going to matter more and more. It's a surprisingly sharp model, especially on long-horizon tasks.
Getting sick of setting up third-party services So I built a skill for it /wizard builds you an interactive CLI for the task you're currently doing, and takes as much work off your hands as possible #1 is how the agent described the wizard, #2-3 is what it looks like:
Sonnet 5 medium is better than GLM 5.2 high and roughly the same price hilarious tbh
Claude Sonnet 5 is now available in Cursor. On CursorBench, it's a meaningful step up from Sonnet 4.6: 57% vs. 49%.
Introducing Nano Banana 2 Lite 🍌 and Gemini Omni Flash 🔮, our new generative media models in the Gemini API and AI Studio! Nano Banana 2 Lite is extremely fast (<4s image) & cheap ($0.034 / 1K image). Omni Flash is SOTA at video editing at $0.10 / sec, same as Veo 3.1 Fast!
Pro-tip: use a token budget when using /goal in the Codex CLI This is an experimental feature we're exploring, you might have to enable the flag. Ask Codex to do it for you.
One unexpected outcome of this is that I'm now using the wiki as the ONLY place I run Claude Code I use it as a master controller for all of my repos, kicking off cross-repo tasks and using Todoist as a task tracker Interesting
Doing my first ever experiments with a personal, entirely agent-managed Karpathy-style wiki X, Discord, Gmail are all being ingested into it every few hours This is the knowledge base that will serve as the environment for all of my future loops
Adding a set of Martin Fowler's code smells from Refactoring into my /review skill Mysterious Name, Duplicated Code, Feature Envy, Data Clumps, Primitive Obsession, Repeated Switches, Shotgun Surgery, Divergent Change... This stuff is catnip for LLMs github.com/mattpocock/ski…
Most agent memory stores text. pantry stores runnable recipes: - your agent fetches the exact saved code by name instead of re-deriving it, - then runs it in its own isolate. pantry never runs your code. @Cloudflare Workers + D1. pantry.coey.dev
If any of the top contributor(s) of openclaw have an idea for a for-profit angle that makes the product EASIER to use than Cowork I'm all ears There needs to be a @wordpressdotcom VIP for OpenClaw that is EXPENSIVE, EASY and has an ABSURD SLA
OpenClaw is now on iOS + Android 🦞 📱 Native mobile apps, finally 💬 Agents in your pocket 🔔 Channels, tasks, replies on the go Run agents from wherever your thumbs are. iOS: apps.apple.com/us/app/opencla… Android: play.google.com/store/apps/det…
Coding agents have made me incredibly Cloudflare pilled So many cool and useful things I never would've bothered with for one off internal tools b/c it was just such a pain in the ass to configure, are now completely free to setup it's amazing
If you're running multiple coding agents (Codex, Claude Code, Pi) at once, you're going to want to watch this. Herdr is a terminal multiplexer that actually understands your agents. Oh, and it has incredible mouse support. Bet I can convince you in 6 minutes:
For anyone hoping that all these open weight models are going to let them launch compete against these "moatless" vibe coding platforms, please enjoy my expensive journey trying to pick the most capable prototyping model x.com/clairevo/statu…
Voice agents now on Vercel
Voice agents, now on Vercel. Realtime, speech and transcription are now live on AI Gateway. Build with 𝚞𝚜𝚎𝚁𝚎𝚊𝚕𝚝𝚒𝚖𝚎, 𝚐𝚎𝚗𝚎𝚛𝚊𝚝𝚎𝚂𝚙𝚎𝚎𝚌𝚑 & 𝚝𝚛𝚊𝚗𝚜𝚌𝚛𝚒𝚋𝚎 on AI SDK 7.
Every time you think you need a dashboard to look at data, stop yourself. Do this instead: 1. Ask your agent to make sure that you have all the data to analyze something actually stored in the database. 2. Ask your agent to write a skill to gather that data. 3. Ask your Show more
code mode, what you need to know: if you give models 10,000 tools in a normal way, it will flood their context window so you need some form of lazy loading you could implement lazy loading to fix this (give the model two tools, `search` and `call`) but the model is still Show more
need someone to explain how code mode works in practice. I understand what its for and the mechanisms but are you dependent on the mcp author to implement it? can i add it in front of a poorly implemented mcp? how can i implement it in my own mcp?
Starting to see more of this actually. Everybody who uses Codex is building some kind of infinitely verifiable evidence receipt evaluator. With Opus 4.6 it was those situation monitoring sites. I can tell what agent you're using by what kind of manic overbuilt software you made
quick example using Think's subagents to drive claude code on cloudflare sandbox. A really nice feature: hijacking egress to send it to cloudflare ai gateway via the binding, so there's ZERO api keys here. I'll expand this to use ai sdk's new harness agent (and make a Show more
I’m giving a talk this week about why I have my coding agents quiz me on code changes, and how I build micro-worlds to understand what’s going on. What questions / topics should I cover?
Mythos / Sol cybersecurity capabilities are equally useful in an offensive as well a defensive capacity. If adversaries get ahold of an equivalent offensive capability, it poses a serious threat to US companies that remain unaware of latent vulnerabilities. In the meantime, I Show more
JUST IN: A new Chinese AI model from Zhipu AI reportedly matches Claude Mythos’ performance at finding security bugs.
in an effort to save token spend, i’ve started asking llms to assign a reasoning level to each task and prompt me to manually switch models whenever it changes… really wish this was possible in codex w/o HITL. is there some config magic i haven’t learnt yet?
"Move non-determinism to the edges, and determinism to the core" - @DavidKPiano Really enjoyed this talk. David shares some thoughts about making programs call LLMs vs have the LLM do everything for us... I've recently been building agents with @flueai and it really resonates Show more
two skills that i love using if you use codex, press cmd+cmd ( left and right cmd buttons at the same time) and just say "make these two skills"
You can now trace and debug eve agent runs with Vercel Observability. Inspect model and tool calls, runtime errors, and token usage in one place. vercel.com/changelog/eve-…
this take is stuck in 2025. codemode solves most of the issues outlined and also discounts nice properties of mcp like elicitation. how does your agent learn about all the clis it has available? skills? their descriptions will do the same to context as poorly designed mcps Show more
please stop giving your agents MCPs and give them CLIs instead 😭🙏 you can load up an agent with 100 cli tools but you can’t load 5 mcps before it shits the bed with context rot all models at this stage are already trained to master terminal commands literally give them cli
Alchemy's Effect IaC abstractions are starting to compound. I integrated Vercel's AI SDK 7 into Alchemy CF Containers so you can deploy harnesses like OpenCode, Claude Code, Codex. Here's OpenCode in a Cloudflare Containers as an importable Resource Layer.
Building high-quality evals is an increasingly important skill. Especially if you're trying to land a job or get into AI, I'd recommend trying to benchmark models on a task/domain you care about. If done well, you'll get the attention of any company training models.
We're sharing new research on how models hack public benchmarks. The latest models, including Opus 4.8 and Composer 2.5, learn to retrieve solutions from the internet or git history. When we apply a stricter harness, eval scores drop significantly.
I guess MCP won. Jokes aside, this is super cool from OpenRouter. Just making it easier for devs to run their long-running agents with the right level of intelligence. More of this, please.
Introducing the OpenRouter MCP, live model intelligence right inside your agent Your agent builds and ships, but when it comes to choosing the right model for the right job, it guesses from 6 month old training data Watch it pick, price, and test the right model:
Work at OpenAI is being transformed by agents, in every department. Across our entire company, people are using Codex to do work that is more complex, longer-running, and increasingly cross-functional. Our internal usage offers an early look at how agentic tools may reshape Show more
ai sdk 7 looks very very good. some breaking changes, so we'll hold off on adding it to agents sdk for a bit while we figure out a nice migration story, and bundle with some other dep changes (probably the new mcp sdk?)
AI SDK 7 is now available. Introducing: reasoning control, agent-level tool approval, tool and runtime context, file and skill uploads, MCP Apps, durable workflows, terminal UI, sandbox support, harness integrations, telemetry, lifecycle events, and more.
Introducing External Agents in Notion: Claude + Cursor. Your team already collaborates in Notion. Now your favorite agents do too. Assign them tasks from a board shared with your whole team. @-mention them like teammates. Watch them run.
Take a look at your favourite skill. Go on, take a look. Check for lines like: - "Make the commit message very detailed" - "Be thorough" - "Make the implementation easy to read" What do these lines have in common? They're no-ops. They do nothing to change the agent's Show more
Executor is joining the YC S26 batch! We're building an open source MCP gateway to connect any agent to any service Your team is constantly spinning up new agents, trying out new tools, wrangling multiple accounts. You need one place to configure everything once, and use them Show more
how to improve your /goal by designing the loop before you run it a poorly designed loop burns tokens and hands you slop fast, that´s why it's important to spend time designing the loop harness instead of writing a /goal and hoping it works, you run it through LOOPER skill Show more
Feeling a bit loopy? Introducing Looper - your loop design coach It’s been made clear that we shouldn’t be promoting agents, we should be designing loops You can use /loop or /goal but if your loop is poorly designed then it won’t produce good output Even worse, a poorly
The most interesting work I've seen to fight AI slop. Design Crit (Criteria-Resolved Image Taste) lets you build a decision layer for vibe/generative design. ie: teach AI to judge design like a designer. Ten professional designers ranked four frontier image models across Show more
Heard good things about Cursor's Thermo-Nuclear Code Quality Review skill, just gave it a shot and can confirm. github.com/cursor/plugins…
Blume.codes 1.0.47 is live! Better agentic codebases🌸 Blume reads your agent conversations locally, then promotes steering and pain signals into code to make the next run better and cheaper. Stop repeating yourself to coding agents - Live on Mac, Linux and Windows
Sakana Fugu Ultra is live on AI Gateway. Mythos-class intelligence in a single call, with a whole pool of models behind it. 𝚖𝚘𝚍𝚎𝚕: '𝚜𝚊𝚔𝚊𝚗𝚊/𝚏𝚞𝚐𝚞-𝚞𝚕𝚝𝚛𝚊' vercel.com/changelog/saka…
Introducing Sakana Fugu: A full multi-agent orchestration system accessible via a single model API. Our ‘Fugu Ultra’ model matches the performance of Fable and Mythos, delivering frontier capability without the risk of export controls. Try it: sakana.ai/fugu 🐡 Show more
The best agent loops need the right tools → agent-browser.dev Verify changes in a real browser → portless.sh No port conflicts. Worktree-friendly. → emulate.dev Emulate third-party APIs → ai-cli.dev Image + video gen via CLI
I don't even prompt/speak to agents that much anymore. With loops, agents do most of it for me now. I do spend more time writing verifiers to provide additional rich instructions (text+audio+images) that help fill in gaps. What's next? Hard to tell!
Bro it’s June 2026. Stop hand editing your prompts. Hold down the dictation button and ramble for 10 minutes. Give the model every fragment, caveat, example, and vibe in your head. It is literally a large language model. If it’s superhuman at anything, it’s reconstructing latent
I wrote a good skill that I think people will like so I'm publishing it but I realized it's really just a template so what i need is a skill to build skills so I had codex go read @mattpocockuk's skills and his skill for building skills, so that my agent can write a skill for Show more
Pretty remarkable what’s happening with open weights AI right now. We’re seeing models achieve SOTA results on specific tasks, and getting close to frontier on some areas of coding and other domains. The more that open weights is able to maintain only a marginal gap from the Show more
I CANNOT believe im saying this right now... but GLM 5.2 in open code is SHITTING on opus 4.8 in claude code. 🤯 how is this possible??
you should have a linter Hands down You should have detailed rules, you should push determinism as far as it can go Use ast analysis to tell your coding agents what needs to be fixed You should absolutely do this BUT If your anti-slop strategy is an LLM and a handful of Show more
I’ve been building my own software factory over the past year. It again changed the way how I do software engineering fundamentally - similar delta compared to coding by hand vs using agents. I’m thinking of writing a blog post about it. What are your biggest questions?
stop writing loops, start writing *control loops* - read current state - read desired end state - one incremental change - repeat* simplest example is a thermostat, but for your code base there’s a reason why some of the best AI coders I know (doing Ralph-style work since Jan Show more
If you're curious how I managed to do over $20,000 in inference on the last 48 hours, here's a video all about it. Spoiler: loops are really powerful
Everyone is using the phrase "one shot" incorrectly, including me! One-shot means that you give the LLM one example before asking it to do the task We mean zero-shot: you give the model a description of what you want done, and no examples
Announcing mattpocock/skills v1 - Achieved a 63% reduction in token cost for skill descriptions - Split skills into model-invocable and user-invocable skills, adding /codebase-design, /domain-modeling, and /grilling - (UPDATED) /writing-great-skills - rewritten from the ground Show more
Introducing eve, an agent framework. 𝚊𝚐𝚎𝚗𝚝/ 𝚊𝚐𝚎𝚗𝚝.𝚝𝚜 𝚒𝚗𝚜𝚝𝚛𝚞𝚌𝚝𝚒𝚘𝚗𝚜.𝚖𝚍 𝚝𝚘𝚘𝚕𝚜/ 𝚜𝚔𝚒𝚕𝚕𝚜/ 𝚜𝚊𝚗𝚍𝚋𝚘𝚡/ 𝚜𝚌𝚑𝚎𝚍𝚞𝚕𝚎𝚜/ Like Next.js, for agents. vercel.com/blog/introduci…
there's a new word i'm hearing a lot in the most frontier-pushingest coding-agent builders: _program design_ for even the best agentic coders trying to maintain code quality, we've all seen it - you come up with something to build - you research the codebase, riff with the Show more
Today we're launching Exa Agent: Opus/GPT 5.5 quality web research at 2-10x lower cost. It's our most powerful endpoint, particularly good at deep research and list-building - as cheaply as possible. Basically you can now use a deep research API for close to the cost of a Show more
Introducing Exa Agent: frontier web research at less than half the cost of GPT 5.5 and Opus. /agent orchestrates a mixture of cost-effective models to complete any web research task, from simple data enrichments to building gigantic lists.
We've gone really quickly from "local models are dogshit" to "local models are good actually" (like, a 12 month window from A to B). I don't think they're actually good ENOUGH yet. We need an Opus 4.5 quality local model. When that happens, I think the world will spill over. Show more
You can now enable design linting in your Codex, Claude Code and Cursor. Yes, really. Install or update Impeccable for glorious *automatic slop and design system drift prevention*. This is not a drill. It dramatically improved the frontend dev performance of all harnesses in Show more
Impeccable 3.7 brings linting to design. Until now it was a skill you asked for help. Now it's a design-system-aware feedback loop that runs while your agent builds, catching slop and design drift before they land. 🪝 Design hooks for Claude, Codex, and Cursor They run after
Claude Code creator: "100% of our pull requests at Anrtopic are run by Claude Code. 80–90% of code review too. The feature I’m using the most today is /loops. I’m not prompting Claude anymore - I’m building loops" in 1-hour interview, Boris reveals his setup, which helps him Show more
Had Hermes Agent with the Manim Video skill plus it's TTS tool create a video explaining Hermes' Agent.
I hate to break it to you guys since we've been trashing loops so much but recursive agent fanout w/ fan for greenfield into /goal or /autoresearch loop with 5.5 xhigh probably is the most effective workflow for raw accuracy and code quality 100% unattended right now
Kimi-K2.7-Code is the new Opensouece SoTA for Coding & Agentic workflows
I feel like there are two secrets to being really productive with ai 1. Understand your problem well enough to describe and design the kinds of tests that will truly catch failures 2. Get your agent to delegate all its work to other agents in clean context windows
Check out Omnigent, an open source harness that lets you use all the existing code harnesses (Claude Code, Codex, OpenCode, pi), collaborate and share sessions in many modalities (e.g. Slack/Teams, cli, webui), while having a fine grained security model that really tightens the Show more
This is number 4: a unified abstraction for all agent harnesses. Write the agent once. Run it on any harness without vendor lock-in vercel.com/changelog/prog… Huge thanks to @felixarntz for pushing this over the line!
My AI haul this year: 1. just-bash 2. Chat SDK 3. deepsec 4. coming soon[*] [*] (4) is "large software" like Chat SDK and so I needed to find a long-term owner before being able to ship it and I just did 🥳
This seems like it's exactly the right shape for coding agents. Not “agent in the editor, good luck” Issue gets filed. Agent investigates. Receipts stay with the task. PR shows up in the same workflow. Human reviews the actual handoff. boring made exciting bc of results and Show more
Linear Agent can now write code. We put it to work, automatically fixing bugs as they land in triage. Igor, the engineer behind these automations, shares how he set them up. linear.app/now/linear-age…
one way to think on loops is that your agent should bring context to you, rather than you bringing context to your agent
We just shipped 𝙷𝚊𝚛𝚗𝚎𝚜𝚜𝙰𝚐𝚎𝚗𝚝, a unified abstraction to orchestrate and integrate any agent’s “brain” into your app. @aisdk now frees you from both model and agent lock-in. (And it doesn’t just get you portability, it’s also delightful to use ofc!)
AI SDK now supports agent harnesses like Claude Code, Codex, and Pi with sandboxed sessions and AI SDK-compatible streams: 𝚌𝚘𝚗𝚜𝚝 𝚊𝚐𝚎𝚗𝚝 = 𝚗𝚎𝚠 𝙷𝚊𝚛𝚗𝚎𝚜𝚜𝙰𝚐𝚎𝚗𝚝({ 𝚑𝚊𝚛𝚗𝚎𝚜𝚜: 𝚌𝚕𝚊𝚞𝚍𝚎𝙲𝚘𝚍𝚎, 𝚜𝚊𝚗𝚍𝚋𝚘𝚡:
AI SDK now supports agent harnesses like Claude Code, Codex, and Pi with sandboxed sessions and AI SDK-compatible streams: 𝚌𝚘𝚗𝚜𝚝 𝚊𝚐𝚎𝚗𝚝 = 𝚗𝚎𝚠 𝙷𝚊𝚛𝚗𝚎𝚜𝚜𝙰𝚐𝚎𝚗𝚝({ 𝚑𝚊𝚛𝚗𝚎𝚜𝚜: 𝚌𝚕𝚊𝚞𝚍𝚎𝙲𝚘𝚍𝚎, 𝚜𝚊𝚗𝚍𝚋𝚘𝚡: Show more
Fable feels superhuman at working over long agentic conversations, sometimes to the point where I can't keep up with what it's telling me 😅 This prompt snippet has been the best fix I've found for getting it to write clearly and drop any jargon:
How I use Claude Code and Remotion to make animated diagrams. Sorry, it's not a single prompt. 1. Find an input language the model knows well. For example, Mermaid for flowcharts. Claude writes it fluently, so it's my entry point. 2. Use Claude to build components that take Show more
Your AI agent can now do your most hated finance tasks. Mercury Skills just launched for Mercury CLI — real, installable AI workflows that live in your terminal.
[AINews] Loopcraft: The Art of Stacking Loops @RichardSSutton has his “Bitter Lesson” for models. We now have the Salty Lesson for agents: Don’t fix things yourself, as you have done historically. Instead focus on systems that scale with more agents, like goals and Show more
Here’s your monthly reminder that you shouldn’t be prompting coding agents anymore. You should be designing loops that prompt your agents.
Anthropic Managed Agents team: "Fable 5 is our best model for running self-improving agent systems. Add /loops, dynamic workflows, dreaming and you are unstoppable" in 13-minutes, Anthropic team shows how to build self-improving agent systems with Fable 5 from scratch. Worth Show more
Here's a simple loop: Tell codex to maintain your repos, wake up every 5 minutes and direct work to threads. That makes it easy to parallelize+steer work as needed. I use a orchestrator skill combined with my triage+autoreview+computer use skills, so some work can land Show more
agent product smell test: 1. makes a slide = toy 2. fills a form = feature 3. checks the form against source docs = useful 4. sends the form, handles the rejection, updates the system = company half of “agentic” is just autocomplete wearing allbirds.
The more I work with Claude Managed Agents, the more I fall in love with this pattern: - Maintain your agent, environment, memory store, vault etc settings with yml and keep them up to date with the ant CLI - Store IDs for all of the objects in a json file - Interact with Show more
we did something similar on cloudflare we have these internal apps that use cf primitives like workers, sqlite, r2 and they're all fronted by cloudflare access which requires SSO 100% vibed by opencode
Everyone's talking about AI-generated HTML. But have you tried giving your sites a zero-config API for saving data, file storage, AI, websockets, etc? We did this at Shopify. Runs on a single VM that costs $200/month, and it's changed the way we work. We call it Quick 👇🧵
fable 5 did a +5k/-5k refactor on our oldest messiest react code, detangling a lot of stuff. Overall I'm very pleased and will be doing more of this, and in more parts of the stack. The diff appears to have no regressions, but still exploring w/ a mix of manual testing and Show more
Everyone's talking about AI-generated HTML. But have you tried giving your sites a zero-config API for saving data, file storage, AI, websockets, etc? We did this at Shopify. Runs on a single VM that costs $200/month, and it's changed the way we work. We call it Quick 👇🧵
new update in @every's Frontier Map: i moved "Closed-loop" product repair closer to productized. Fable makes it newly possible to speed-run through backlogs of issues and close bugs and paper cuts almost as soon as they're opened. @kieranklaassen is doing this with Show more
wondering why I feel exhausted. maybe: the agents do all the easy stuff, and I have to work through the leftover hard bits, which means I'm perpetually locked in. and as the models get better, "my" work just gets harder and harder, until I'm basically underqualified to do the Show more
the future of all work is this. You must define: - a goal - the criteria that define it - the verifier that makes sure it is achieved - the sensors that inform the verifier - the actuators that affect the sensors - The envelope that contains the sensors and actuators
The codex "goal" feature is a really good way to spend dozens of hours optimizing some total bullshit btw. If your final criteria is it all vague it will specification game and make masturbatory "evidence" and "verifiers" and "gates" and "smoke tests". must be hell internally
btw Fable 5 in Claude Code with no system prompt (claude --system-prompt ".") is friend shaped
Did you know Codex threads can now manage other Codex threads? You can use this to make a chief of staff merge-bot with a pretty simple prompt. Tired: I use a thread to implement some change, put up the PR, babysit CI for hours, and the agent often fails over to me to route from Show more
We've reset 5-hour and weekly rate limits for all users. Enjoy Fable 5!
fuck loops, use queues or you are dead. literally dead or worse, irrelevant. tomorrow we are using ring buffers.
We talk a lot about how important it is to set up self-verification loops. Especially in the age of powerful models that can run for long periods of time, self-verification is a key ingredient that enables the model to run for much longer, delivering a result that is closer to Show more
How do you get Claude Code to check its own work before handing it back? Watch how you can encode your manual checks so Claude closes its own feedback loop:
tool calls are token hogs. if you can make a single tool call that then invokes 5 api requests rather than 5 tool calls it adds up. search() + execute() are the only tools you need as it turns out blog.cloudflare.com/code-mode-mcp/
code mode has helped @cloudflare save ~93% on token cost talking to peers in the industry whose spend is creeping towards upwards of $100k per engineer per year you do the math on the cost savings of that for a large engineering org
in the past few days the Cloudflare MCP server made 2.6M API requests from 735k agent tool calls. code mode enables agents to do more per tool call.
I scraped every public Claude Code dynamic workflow: 1,245 across 500+ repos. The link is below and here are some findings after ranking and classifying them. Workflow is a JS script that orchestrates a fleet of subagents. The plan lives in code (loops, branches, fan-out) and Show more
It’s good framing. Queues encourages you to break down the problem and asyncify one part of the process at a time I see three failure modes for agent coding loops: 1) people who spend all their time building the thing that builds the thing and don’t see results (or Show more
Your issue tracker is a queue of tasks Agents pick tasks off that queue, and complete them They then go into a different queue - human review The cycle begins anew
Everyone's banging on about loops When they should be thinking about queues
Just landed nested subagent support in Claude Code Starting to experiment more with agents kicking off agents as a way to better manage context. Capped at depth=5 to start, going out in today’s release. Lmk what you think!
I know folks are side eyeing the loop discourse but this is exactly how I work now. I use scheduled @goose_oss recipes + skills + MCPs + subagents for most of my work. It may be hard to grasp how that translates to coding work. Coding is one part of the overall routine. Show more
Peter and Boris's “loops, not prompts” point is getting backlash because people are reacting to the slogan but missing the direction of the field. I think “loop” is better framed as self-improving context infrastructure. Agents are not magically improving themselves. We are Show more
Here’s your monthly reminder that you shouldn’t be prompting coding agents anymore. You should be designing loops that prompt your agents.
I want an agent workflow tool. It should let me: - describe work items and plans - assign tasks to agents - review code diffs - work multiplayer with my team - collect customer context from many places -define shared skills and MCP servers - use it fully from slack
I've been really surprised by the number of people I consider cracked engineers who are shitting on this post. I've been building with effective loops in mind for a while, and I would say I've built more effective products in the last 3 weeks than I have in the last year.
Here’s your monthly reminder that you shouldn’t be prompting coding agents anymore. You should be designing loops that prompt your agents.
Are you making a CLI tool? If so, you should think hard about how easy and intuitive it is to use. And not by humans, by agents! Or outsource that work to my skill, which does it all for you in an incredibly intensive, scientific, and comprehensive way: jeffreys-skills.md/skills/agent-e…
Used the 'agent-ergonomics-and-intuitiveness-maximization-for-cli-tools' skill by @doodlestein this morning. Honestly, that skill name could almost be recast as a novel. Anyway, I love it. My budding CLI is now much more agent-friendly.
The best rule I stole from @zeke agents.md: "When corrected, propose an edit to AGENTS.md so the same mistake doesn't recur." Your agent learns from its mistakes permanently. One line that has a huge effect. What are the most impactful rules are in your agents.md?
Skills™ are great, but you can get a lot of mileage from a few precise edits to your global AGENTS.md file. Here are my Cloudflare pointers that fill the knowledge cutoff gap and give my agents some context about my account setup:
When we first demoed Claude Code internally, it got two reactions on Slack. A year after GA, @_catwu and I sat down to talk about what's changed: why I use auto mode instead of plan mode, how routines fix bugs before I see them, why I do most of my coding from my phone now, and Show more
Claude Code's first demo got two Slack reactions. One year after GA, @bcherny and @_catwu look back: verification best practices, why we built auto mode, routines and loops, and what's next. youtube.com/watch?v=Hth_tL…
"Human in the loop" makes the human sound like a safety widget bolted onto a machine. That's the wrong mental model for almost every B2B workflow with real consequences. Cybernetics has a sharper frame: control and communication. It changes which AI features you ship. I made Show more
introducing loops! a directory of pre-built agent workflows for Cursor, Claude Code, and other coding agents. copy a kickoff. set exit conditions. let the agent loop until the job is actually done. 26 loops live → loops.elorm.xyz
Here’s your monthly reminder that you shouldn’t be prompting coding agents anymore. You should be designing loops that prompt your agents.
for a really conceptually simple example of a loop, check out agent-reviews.com a skill + cli combo that continuously polls GitHub PR reviews, fixes bugs, which then triggers re-reviews, and doesn’t stop until every AI code review bot is satisfied.
/teach is live Learn anything, from rubik's cube to vocal harmonies to software fundamentals. npx skills add mattpocock/skills --skill teach Best skill I've ever built, video coming soon github.com/mattpocock/ski…
i don't think "sandboxes" will win as a standalone category it's just too easy to swap providers it's everything around it that will make people stay: orchestration, state, networking, observability, security, deployment workflows, pricing, reliability... sandboxes are a Show more
don't use loops, design state machines
Here’s your monthly reminder that you shouldn’t be prompting coding agents anymore. You should be designing loops that prompt your agents.
The numbers may be a bit extreme here, but unquestionably use-cases have to stratify in the next year or two between model families. We’ll see a split between frontier intelligence for high end tasks and work, and much cheaper models for high volume workloads that can Show more
Good take My guess is - demand for intelligence is near infinite - but 80% of workloads will be running on 99% cheaper models within 12-18 months - 20% of workloads will still run on latest gen models where IQ maxing is important (scientific breakthroughs, higher level
/no-mistakes is here! by popular demand i've made the most impactful tool in my agentic engineering setup "no-mistakes" invocable as a skill in Claude Code, Codex et al just type "/no-mistakes" once your agent has made changes, and watch the magic unfold details below 👇
been thinking about this a lot and grappling with the sheer breadth of the ai coding experience/skill spectrum Every time I see someone say “stop building the thing and instead build the thing that builds the thing” I eyeroll super hard at this engineers-playing-ant-farm slop Show more
"you should be running rabbitmq and piping tasks into it and so your agent can generate max slop while you pretend your output is valuable" im so tired of hearing about "groundbreak techniques" to produce more unmaintainable software
How I made an agent that buys our groceries, runs the family budget, and is planning a move:
the more I work with complex subagent orchestration (dynamic workflows, etc) the stronger my belief that more people on a project ≠ better outcomes
This is getting a lot of hate. But he's correct. Not just for coding, but all work. It's always about the correct abstraction level that the loops run at, and how to verify. Once you see this, and then couple it with techniques like Auto-Researcher, you can see the future.
People are confused about what "loop" means in the context of LLMs. Stop babysitting the model. Build a non-interactive AI application instead. Loops can be simple (ralph, autoresearch) or complex (fabro.sh). Your task: "build the thing that builds the thing."
Spent some time over the weekend building a scraper to collect all existing dynamic workflows. The feature only appeared about a week ago, so there aren't that many in the wild yet - around 800. Based on GitHub activity, roughly 100 new ones are being published every day. Almost Show more
A dynamic workflow like this consumes ~10M tokens but delivers a comprehensive market research report that previously would have required dozens of deep research runs and hours of data compilation I'm running this for every project that applies to @cyberfund and gets through
right now I do this by structuring as much context in a workspace, setting very clear requirements & contracts upfront, and working with an orchestrator agent to launch workflows/goals/loops that all basically do the same thing - intake the context, begin whatever implementation, Show more
Here’s your monthly reminder that you shouldn’t be prompting coding agents anymore. You should be designing loops that prompt your agents.
Here’s your monthly reminder that you shouldn’t be prompting coding agents anymore. You should be designing loops that prompt your agents.
My ideal agent harness: - Works on my iPhone - Connects to GitHub - Runs on a schedule - Sends me UI screenshots with changes it made - Can set system prompts for each project - One long-running chat with a “manager” agent that dispatches subagents for each task - Can use both Show more
this is interesting in at least one anecdotal setting I've seen, our Reactor Harness uses 150x fewer tokens than the raw Claude Code equivalent for the task of keeping a model of my local agent usage up to date this 150x improvement is not on a "per request" basis. instead, Show more
Curiously enough I did office hours today with a startup that cuts companies' LLM token costs by optimizing requests. They can cut costs by about half, which they split with the customer. So the TAM is a quarter of the model companies' corporate revenue. That's a big TAM!
The more I use agents, the more I think the human job is the bread, not the meat. Sweat the idea up front. Let agents do the messy middle. Then actually use the thing and feed back what you learn. The middle is getting cheaper. The bookends are where product judgment Show more
Once you understand the power of generalized agents, you should likely get off of them and build your bare bones minimalist ones yourself. Custom tuned to your preferences. You do not want all of the tech debt slop going into them, and you should be agent agnostic anyways.
I have a similar setup of 3 skills that I now run on almost every feature after I am done coding. 1. Review skill Spins up 3 sub-agents using different models from different providers. They independently look for performance issues, over-engineering, security vulnerabilities, Show more
We doubled Claude Cowork usage limits for the next month. This applies to your 5-hr rate limits. If you’ve been saving up a big messy project, now’s the time.
We've doubled usage limits in Claude Cowork for the next month. Delegate bigger, more complex tasks to Claude.
"Codex Sites" is literally just the Cloudflare plugin in a trenchcoat It solves exactly 1 problem: creating your own Cloudflare account If only there were a protocol to let agents create their own accounts or pay for things... Oh wait! Stripe Projects and x402. I am so Show more
how i hit inbox 0 every day with Codex:
Stumbled upon a Codex skill that creates cool illustrations to explain topics or tell stories. You feed it text (blog, article, narrative, even code) and it makes explainer graphics with this cute blob character. I gave it the repo for the X recommendation algo and got this 👇
I can’t believe people treat hermes/openclaw as a coding agent on par with the major players. I too can prompt a slackbot to make changes and it’s garbage for productivity. There’s a reason harnesses are converging back to IDE design.
Git worktrees are an anti-pattern in agentic software development, even if the Claude Code team recommends them as one of their "top tips"! If you want tips on playing the piano well, you don't go to Steinway & Sons, you ask Evgeny Kissin or Lang Lang! And try my skills site!
I went HAM on git worktrees when I learned they were a thing like 6 months ago but slowly drifted back to single branch flows. It's just way easier to manage and far less repeat/conflicting work. But how do you make sure multiple agents don't collide? The flow is simple: You
we landed on a pretty good workflow for doing parallel work in OpenCode this demo is with git worktrees but i also preview an alternative we're working on at the end this will be in 1.6.0
An app can be a home-cooked meal (2020) personal software was a bit early in 2020 but in 2026, it really can be as personal as a home cooked meal, or a handwritten letter robinsloan.com/notes/home-coo… Show more
Your favourite frontier model claims it has a million token context window but the longer you speak to it, the dumber it gets. It’s called attention dilution, and it’s why your sessions are giving you AI slop. Here is how to fix it in under 2 minutes:
In the first episode of our new series Full Stack, @conductor_build CEO and co-founder @charlieholtz takes us into the details of how he sets up his workflow for coding and managing AI agents. 00:00 – Building Conductor With Conductor 01:05 – Managing a Team of Coding Agents Show more
How I keep my Agent Skills in sync across projects 00:25 Installing skills with npx skills add 02:18 Using your own private skills repo 03:30 Installing from a personal repo 04:35 Pushing local skills back with skill push 05:35 Why compounding skills matters
Anthropic runs hundreds of claude code skills internally they just published everything they learned the biggest misconception people have: skills are "just markdown files" they're not. they're folders. that means scripts, assets, reference code, data, templates all Show more
I’ve been dumping on OpenAI with low effort meme tweets that get too many views, but Codex is the best DevX acceleration product of all time and I wrote about it here: cpojer.net/posts/modern-e…
Excited to share how Anthropic's data team has automated 95% of business analytics queries with Claude. Blog post covers how we approach evals, ablations, and online validation!
How do we automate business analytics with Claude? New blog post covering our best practices for skills, data foundations, and evaluations when building agents to perform data analysis: claude.com/blog/how-anthr…
Just open sourced my variant of this amazing skill that I'm now using a LOT github.com/alexknowshtml/… - discovers its own context for a topic from session files - progress indicators and visual recaps - solo or co-learner modes It rules so hard
been asking others at Anthropic how they stay in the loop with Claude and fully understand the work being done this is one of my favorites from Suzanne:
yesterday I danced to afrobeats in the sun while shipping a core feature for my startup. 5 giant Claudes floated around our terrace in my Vision Pro. I steered them by looking, pointing & speaking at them. you now can buy Tony Stark's sci-fi workshop for ~$2k + a week of setup.
A harnessed LLM agent, clearly explained! Most people picture this as a model with tools bolted on. The real architecture inverts that relationship. The model itself is deliberately thin. Intelligence gets pushed outward, and the harness composes it at runtime. Three Show more
codex has replaced ChatGPT as my daily driver; not b/c I'm coding all day; it helps me move actual work forward: research, analysis, content, ops, planning, and execution. codex for everything.
Exclusive: After falling behind Anthropic in coding, OpenAI built Codex into one of its fastest-growing products and is now making it central to ChatGPT’s future. Full story: thein.fo/3PT3Huu
YES-CODE An entire category of software, "no-code", was built under the presumption that code is expensive, difficult, and scarce. Coding agents have forever changed the equation. Code is now cheap, easy, and abundant. I remember @cramforce being asked by an analyst long ago: Show more
Our warp[dot]dev site gets 10M visitors/year. We migrated the whole thing from a no-code editor back to code in just 3 weeks. Very few hiccups, and SEO actually improved. Plus, the marketing team is free to use Warp to ship future changes
Workflows are the biggest upgrade to Claude Code’s capabilities since skills and subagents. I dove deep into it with @sidbid to figure out best practices, examples and more. I’m particularly excited about the non-technical tasks it enables for Claude Code.
Would it shock anyone if I told you Opus 4.6 performs better than 4.7 and 4.8
If your daughter needs tutoring in algebra, you can probably find someone cheaper than Albert Einstein. Giving every task to GPT5.5 or Opus 4.8 is overkill. Often times you can get the task done just as well, but 10x cheaper and 10x faster with a smaller model.
Introducing model routing to Factory. Factory Router picks the right model for every task, automatically. Maintain frontier performance while cutting costs by 25%.
This is right & is what I've been telling anyone who'll listen (= few). I'd qualify this whole thread even further tho: the step-function bump we saw in coding capability at the end of '25 was a one-off. The models are near the top of the S curve in coding capability.
There are several reasons software was likely disrupted first: 1) The work is already digital and performed through text 2) Coding has tight feedback loops that allow for easy testing of whether the AI output worked 11/n
LLMs still seem bad at making some key "give up and start from scratch" call that seasoned software engineers are better at making. Latest example: I was implementing (w/ Opus 4.8max) a feature to add offline mode to an internal app. I did a PR, it got mired in an endless Show more
The official /codex plugin for Claude Code is pretty sweet. /codex:review to have Codex critique CC's work /codex:adversarial-review for increased spice /codex:rescue to handoff a mess only GPT 5.5 can untangle It's like /handoff with less ceremony. Good in a pinch!
Some tips to help agents understand your codebase: 1. The source code either needs to be the source of truth, or have something legible as a path to the source. For example, if marketing site content is actually stored in a CMS, you need to either delete the CMS and move that Show more
Cursor just got a major upgrade! Agents can onboard to your codebase, use a cloud computer to make changes, and send you a video demo of their finished work. The latency of using the remote desktop is smooooth.
slowly we're all realizing that tools should be called from code, not from within the llm api
Introducing Search as Code, our new search architecture for AI agents. It writes Python that calls our search stack directly, instead of looping through function calls one at a time. Available in the Perplexity Agent API, and now default in Computer. research.perplexity.ai/articles/rethi…
Opus 4.8 just broke ARC-AGI-3 it tripled GPT-5.5's score we are now at a breathtaking 1.5% human efficiency
lot of talk today about "vibe coding" / "is vibe coding dead" - its not. The problem is, lots of us, on various timelines, tried to apply vibe coding to engineering our systems in production and realized that *really* doesn't work it's always been a useful thing, I *also* vibe Show more
using AI for coding is a deeply technical engineering craft most people don't approach it as so, and don't get the results we associate with high craft but the ones who do have been sprinting ahead more tokens wont save you, more thinking + skill + llm intuition will have