Tweets from the frontier

Tuesday, August 18, 2026

Monday, August 17, 2026

Sunday, August 16, 2026

Saturday, August 15, 2026

Friday, August 14, 2026

Thursday, August 13, 2026

Wednesday, August 12, 2026

Tuesday, August 11, 2026

Monday, August 10, 2026

Sunday, August 9, 2026

Saturday, August 8, 2026

Friday, August 7, 2026

Thursday, August 6, 2026

Tuesday, August 4, 2026

Monday, August 3, 2026

  • I've been calling this the "prompting paradox" concept for about a year now. LLMs can solve pretty much any problem you specify well enough, and the entire idea now is to help teach it how to specify things better for itself !

    xjdr
    xjdr
    @_xjdr

    i saw Terrence Tao use sol med to answer a lot of very complex problems in one of his chat logs. i became curious. i had a particularly sticky problem that was in my 'ai cant do this yet' pile that i was only very recently able to get sol ultra to solve correctly (the problem

    129
    Reply

Saturday, August 1, 2026

  • It’s become pretty clear what the next 10yrs are gonna look like: If you want job security: - build a harness, two, three - read x all day, try everything new that gains traction - try out all new sdks / agent frameworks If you want gen wealth: - do the same, but also post Show more

    kache
    kache
    @yacineMTB

    My father in law's engineering office had a letter boy travel from cubicle to cubicle, carrying off spec sheets to the PCB designers. Replaced by email. It's the same thing

    1.1K
    Reply
  • My AI coding journey so far Copy paste ChatGPT to jetbrains Cursor autocomplete Codebuff CLI + jetbrains for reading + cursor for polish Claude code in terminal + jetbrains for debugging 4 Claude’s in tmux worktrees with a 5th merging every commit into main Claude code - Show more

    Greg Kamradt
    Greg Kamradt
    @GregKamradt

    My AI coding journey so far: * Copy paste between ChatGPT & vscode * Cursor autocomplete * Cursor sidebar * Claude code in cursor's terminal * Claude code/Codex in terminal, but no ide * Amp w/ 5 terminals at once * Codex Desktop App * Codex Desktop App + mobile

    252
    Reply

Friday, July 31, 2026

Thursday, July 30, 2026

Wednesday, July 29, 2026

Tuesday, July 28, 2026

Monday, July 27, 2026

Sunday, July 26, 2026

Saturday, July 25, 2026

Friday, July 24, 2026

Thursday, July 23, 2026

Wednesday, July 22, 2026

  • there's basically no difference in the effectiveness of MCP and CLIs + skills for agents, agents are equally effective at both CLIs are fine, but require basically a full sandbox to run making them a non starter for lightweight agents they're also lossy - you can't know without Show more

    Nick Vasilescu
    Nick Vasilescu
    Orgo
    @nickvasiles

    so what's the consensus on CLI + Skills vs MCP for agents? is the difference enough to care about? or just use whatever is available/works?

    239
    Reply

Tuesday, July 21, 2026

Monday, July 20, 2026

  • My theory: Opus 5(.1) was meant to replace Fable 5 for most dev work. It would be cheaper and bench nearly as well. My guess is that it didn’t perform as well as they hoped, and the delays on Fable were attempts to improve the new Opus. I would *guess* we’ll still see a be Opus Show more

    Claude
    Claude
    Anthropic
    @claudeai

    Beginning July 20, Claude Fable 5 will be included in all Max and Team Premium plans, at 50% of limits. Pro and Team Standard users will continue to have access to Fable via usage credits, and will receive a one-time $100 credit. Demand for Fable has been challenging to

    4.2K
    Reply

Sunday, July 19, 2026

Saturday, July 18, 2026

Thursday, July 16, 2026

Wednesday, July 15, 2026

Tuesday, July 14, 2026

Monday, July 13, 2026

Sunday, July 12, 2026

Saturday, July 11, 2026

Friday, July 10, 2026

Thursday, July 9, 2026

Wednesday, July 8, 2026

  • GPT-5.6 is like a AWD hybrid minivan, i have irrational love for her. a modern marvel. reliable. you can technically get it murdered out (dark mode.) it will take you where you need to go. you'll pass it own to your teenager. Fable is like going to the future in a DeLorean

    Dan Shipper 📧
    Dan Shipper 📧
    @danshipper

    GPT-5.6 is like a Porsche, Fable is like a warp drive. We've been testing internally @every for about a month. And GPT-5.6 is the best combination of power, speed, and performance for your day to day knowledge work and coding. Fable is a different beast. If you need to get

    99
    Reply
  • Anthropic has given us an engineering manager and OpenAi has gifted us a top 1% engineer ❤️ Use the two together in your work to get the best of both worlds. Prompt: Tell Claude “install the codex CLI and use it within Claude code as a sub agent. Default to GPT 5.6 Sol (or Show more

    OpenAI
    OpenAI
    @OpenAI

    GPT-5.6 Sol, along with Terra and Luna, will launch publicly this Thursday. We’re expanding preview access globally now.

    Image
    334
    Reply
  • i have this file to replace the claude code system prompt with a pi-style one alias claude-pi='claude --system-prompt-file ~/.claude/pi-system-prompt.txt --dangerously-skip-permissions'

    Image
    Matt Pocock
    Matt Pocock
    @mattpocockuk

    Here's a step-by-step process to kill all the bloat from your Claude Code system prompt: 1. Run a proxy so you can see exactly what gets sent to Claude Code (included in the article) 2. "Fuck, there is so much cruft in there" 3. Use my settings.json to kill all the bloat Down

    31
    Reply

Tuesday, July 7, 2026

Monday, July 6, 2026

Sunday, July 5, 2026

Saturday, July 4, 2026

Friday, July 3, 2026

Thursday, July 2, 2026

  • eve is like Next.js, for agents framework for building agents. one folder, durable by default ↓ x.com/vercel/status/…

    Vercel
    Vercel
    @vercel

    Introducing eve, an agent framework. 𝚊𝚐𝚎𝚗𝚝/ 𝚊𝚐𝚎𝚗𝚝.𝚝𝚜 𝚒𝚗𝚜𝚝𝚛𝚞𝚌𝚝𝚒𝚘𝚗𝚜.𝚖𝚍 𝚝𝚘𝚘𝚕𝚜/ 𝚜𝚔𝚒𝚕𝚕𝚜/ 𝚜𝚊𝚗𝚍𝚋𝚘𝚡/ 𝚜𝚌𝚑𝚎𝚍𝚞𝚕𝚎𝚜/ Like Next.js, for agents. vercel.com/blog/introduci…

    296
    Reply
  • One powerful pairing of artifacts 🤝 loops is to have Claude keep an artifact up to date with a periodic report. E.g. scan these results every day and update this artifact with X, Y, Z metrics Great little way to build shareable, short-lived (several weeks) dashboards.

    ClaudeDevs
    ClaudeDevs
    Anthropic
    @ClaudeDevs

    Artifacts in Claude Code are now also available on Pro and Max plans. Ask for an artifact, Claude writes the code, publishes it live to claude.a‍i, and updates it in real time while it keeps working. Pages are private to your account and fully self-contained.

    9
    Reply
  • I really like this use case. Giving agents high fidelity context is key. Words can only go so far. Code in Notion is just getting started! The team is firing on all cylinders. 💥

    Notion
    Notion
    @NotionHQ

    Turn a PRD into something people can poke at. Your agent can take the doc, sketch the flow, and build a lightweight prototype on the page. Useful enough for feedback before anyone opens Figma.

    96
    Reply
  • Artifacts in Claude Code are now also available on Pro and Max plans. Ask for an artifact, Claude writes the code, publishes it live to claude.a‍i, and updates it in real time while it keeps working. Pages are private to your account and fully self-contained.

    Claude
    Claude
    Anthropic
    @claudeai

    New in Claude Code: Artifacts. Interactive pages built from your session, like a PR walkthrough or a living project dashboard, shared with your team at a private link. Available in beta on Team and Enterprise plans.

    8.5K
    Reply
  • I absolutely adore this whole talk-turned-tweet-thread – go read the whole thing! – but i want to call-out this demo in particular We now assume that AI will do the work in a loop while we do something else – *and then when it's done* we may try to understand it However, as we Show more

    Geoffrey Litt
    Geoffrey Litt
    Notion
    @geoffreylitt

    Another example. I was migrating my personal website from one framework to another, and Claude wrote a script that did it — something like this. But it was very hard to review: I wasn't familiar with the new framework, and all I could say was "I guess that looks about right." So

    80
    Reply
  • This is part of the premise of Kody 🐨 The goal is an intentional pipeline that converts your agent-manual processes into deterministic code as much as possible to improve performance and reliability as well as reduce costs.

    Ben Vinegar
    Ben Vinegar
    Modem
    @bentlegen

    I am taking the opposite position: As code becomes more plentiful and more opaque, well-architected and well-tested libraries matter as much as they ever did.

    24
    Reply

Wednesday, July 1, 2026

Tuesday, June 30, 2026

Monday, June 29, 2026

Sunday, June 28, 2026

  • code mode, what you need to know: if you give models 10,000 tools in a normal way, it will flood their context window so you need some form of lazy loading you could implement lazy loading to fix this (give the model two tools, `search` and `call`) but the model is still Show more

    Oliver
    Oliver
    Mesa
    @olvrgln

    need someone to explain how code mode works in practice. I understand what its for and the mechanisms but are you dependent on the mcp author to implement it? can i add it in front of a poorly implemented mcp? how can i implement it in my own mcp?

    348
    Reply
  • Starting to see more of this actually. Everybody who uses Codex is building some kind of infinitely verifiable evidence receipt evaluator. With Opus 4.6 it was those situation monitoring sites. I can tell what agent you're using by what kind of manic overbuilt software you made

    rUv
    rUv
    @rUv

    Anything you can build, I can build better. #ycombinator #demoday

    Image
    397
    Reply

Saturday, June 27, 2026

Friday, June 26, 2026

Thursday, June 25, 2026

  • Alchemy's Effect IaC abstractions are starting to compound. I integrated Vercel's AI SDK 7 into Alchemy CF Containers so you can deploy harnesses like OpenCode, Claude Code, Codex. Here's OpenCode in a Cloudflare Containers as an importable Resource Layer.

    Image
    90
    Reply
  • Building high-quality evals is an increasingly important skill. Especially if you're trying to land a job or get into AI, I'd recommend trying to benchmark models on a task/domain you care about. If done well, you'll get the attention of any company training models.

    Cursor
    Cursor
    SpaceXAI
    @cursor_ai

    We're sharing new research on how models hack public benchmarks. The latest models, including Opus 4.8 and Composer 2.5, learn to retrieve solutions from the internet or git history. When we apply a stricter harness, eval scores drop significantly.

    Image
    1.6K
    Reply
  • I guess MCP won. Jokes aside, this is super cool from OpenRouter. Just making it easier for devs to run their long-running agents with the right level of intelligence. More of this, please.

    OpenRouter
    OpenRouter
    @OpenRouter

    Introducing the OpenRouter MCP, live model intelligence right inside your agent Your agent builds and ships, but when it comes to choosing the right model for the right job, it guesses from 6 month old training data Watch it pick, price, and test the right model:

    275
    Reply
  • ai sdk 7 looks very very good. some breaking changes, so we'll hold off on adding it to agents sdk for a bit while we figure out a nice migration story, and bundle with some other dep changes (probably the new mcp sdk?)

    AI SDK
    AI SDK
    Vercel
    @aisdk

    AI SDK 7 is now available. Introducing: reasoning control, agent-level tool approval, tool and runtime context, file and skill uploads, MCP Apps, durable workflows, terminal UI, sandbox support, harness integrations, telemetry, lifecycle events, and more.

    AI SDK 7
Develop, run, and observe agents.
    156
    Reply

Wednesday, June 24, 2026

Tuesday, June 23, 2026

  • how to improve your /goal by designing the loop before you run it a poorly designed loop burns tokens and hands you slop fast, that´s why it's important to spend time designing the loop harness instead of writing a /goal and hoping it works, you run it through LOOPER skill Show more

    Image
    Kevin Simback 🍷
    Kevin Simback 🍷
    @KSimback

    Feeling a bit loopy? Introducing Looper - your loop design coach It’s been made clear that we shouldn’t be promoting agents, we should be designing loops You can use /loop or /goal but if your loop is poorly designed then it won’t produce good output Even worse, a poorly

    1.7K
    Reply

Monday, June 22, 2026

Sunday, June 21, 2026

Saturday, June 20, 2026

Friday, June 19, 2026

Thursday, June 18, 2026

Wednesday, June 17, 2026

Tuesday, June 16, 2026

  • Today we're launching Exa Agent: Opus/GPT 5.5 quality web research at 2-10x lower cost. It's our most powerful endpoint, particularly good at deep research and list-building - as cheaply as possible. Basically you can now use a deep research API for close to the cost of a Show more

    Image
    Image
    Exa
    Exa
    @ExaAILabs

    Introducing Exa Agent: frontier web research at less than half the cost of GPT 5.5 and Opus. /agent orchestrates a mixture of cost-effective models to complete any web research task, from simple data enrichments to building gigantic lists.

    242
    Reply
  • You can now enable design linting in your Codex, Claude Code and Cursor. Yes, really. Install or update Impeccable for glorious *automatic slop and design system drift prevention*. This is not a drill. It dramatically improved the frontend dev performance of all harnesses in Show more

    Impeccable
    Impeccable
    @impeccable_ai

    Impeccable 3.7 brings linting to design. Until now it was a skill you asked for help. Now it's a design-system-aware feedback loop that runs while your agent builds, catching slop and design drift before they land. 🪝 Design hooks for Claude, Codex, and Cursor They run after

    656
    Reply

Sunday, June 14, 2026

Saturday, June 13, 2026

Friday, June 12, 2026

  • This seems like it's exactly the right shape for coding agents. Not “agent in the editor, good luck” Issue gets filed. Agent investigates. Receipts stay with the task. PR shows up in the same workflow. Human reviews the actual handoff. boring made exciting bc of results and Show more

    Linear
    Linear
    @linear

    Linear Agent can now write code. We put it to work, automatically fixing bugs as they land in triage. Igor, the engineer behind these automations, shares how he set them up. linear.app/now/linear-age…

    17
    Reply
  • We just shipped 𝙷𝚊𝚛𝚗𝚎𝚜𝚜𝙰𝚐𝚎𝚗𝚝, a unified abstraction to orchestrate and integrate any agent’s “brain” into your app. @aisdk now frees you from both model and agent lock-in. (And it doesn’t just get you portability, it’s also delightful to use ofc!)

    Vercel Developers
    Vercel Developers
    Vercel
    @vercel_dev

    AI SDK now supports agent harnesses like Claude Code, Codex, and Pi with sandboxed sessions and AI SDK-compatible streams: 𝚌𝚘𝚗𝚜𝚝 𝚊𝚐𝚎𝚗𝚝 = 𝚗𝚎𝚠 𝙷𝚊𝚛𝚗𝚎𝚜𝚜𝙰𝚐𝚎𝚗𝚝({ 𝚑𝚊𝚛𝚗𝚎𝚜𝚜: 𝚌𝚕𝚊𝚞𝚍𝚎𝙲𝚘𝚍𝚎, 𝚜𝚊𝚗𝚍𝚋𝚘𝚡:

    1.2K
    Reply
  • [AINews] Loopcraft: The Art of Stacking Loops @RichardSSutton has his “Bitter Lesson” for models. We now have the Salty Lesson for agents: Don’t fix things yourself, as you have done historically. Instead focus on systems that scale with more agents, like goals and Show more

    Image
    Peter Steinberger 🦞
    Peter Steinberger 🦞
    OpenClaw🦞
    @steipete

    Here’s your monthly reminder that you shouldn’t be prompting coding agents anymore. You should be designing loops that prompt your agents.

    149
    Reply

Thursday, June 11, 2026

Wednesday, June 10, 2026

  • we did something similar on cloudflare we have these internal apps that use cf primitives like workers, sqlite, r2 and they're all fronted by cloudflare access which requires SSO 100% vibed by opencode

    Daniel Beauchamp
    Daniel Beauchamp
    Shopify
    @pushmatrix

    Everyone's talking about AI-generated HTML. But have you tried giving your sites a zero-config API for saving data, file storage, AI, websockets, etc? We did this at Shopify. Runs on a single VM that costs $200/month, and it's changed the way we work. We call it Quick 👇🧵

    1.2K
    Reply
  • the future of all work is this. You must define: - a goal - the criteria that define it - the verifier that makes sure it is achieved - the sensors that inform the verifier - the actuators that affect the sensors - The envelope that contains the sensors and actuators

    🎭
    🎭
    @deepfates

    The codex "goal" feature is a really good way to spend dozens of hours optimizing some total bullshit btw. If your final criteria is it all vague it will specification game and make masturbatory "evidence" and "verifiers" and "gates" and "smoke tests". must be hell internally

    805
    Reply

Tuesday, June 9, 2026

Monday, June 8, 2026

  • Peter and Boris's “loops, not prompts” point is getting backlash because people are reacting to the slogan but missing the direction of the field. I think “loop” is better framed as self-improving context infrastructure. Agents are not magically improving themselves. We are Show more

    Peter Steinberger 🦞
    Peter Steinberger 🦞
    OpenClaw🦞
    @steipete

    Here’s your monthly reminder that you shouldn’t be prompting coding agents anymore. You should be designing loops that prompt your agents.

    86
    Reply
  • I've been really surprised by the number of people I consider cracked engineers who are shitting on this post. I've been building with effective loops in mind for a while, and I would say I've built more effective products in the last 3 weeks than I have in the last year.

    Peter Steinberger 🦞
    Peter Steinberger 🦞
    OpenClaw🦞
    @steipete

    Here’s your monthly reminder that you shouldn’t be prompting coding agents anymore. You should be designing loops that prompt your agents.

    46
    Reply
  • Are you making a CLI tool? If so, you should think hard about how easy and intuitive it is to use. And not by humans, by agents! Or outsource that work to my skill, which does it all for you in an incredibly intensive, scientific, and comprehensive way: jeffreys-skills.md/skills/agent-e…

    Image
    James Johnson
    James Johnson
    @JJ90_James

    Used the 'agent-ergonomics-and-intuitiveness-maximization-for-cli-tools' skill by @doodlestein this morning. Honestly, that skill name could almost be recast as a novel. Anyway, I love it. My budding CLI is now much more agent-friendly.

    72
    Reply
  • The best rule I stole from @zeke agents.md: "When corrected, propose an edit to AGENTS.md so the same mistake doesn't recur." Your agent learns from its mistakes permanently. One line that has a huge effect. What are the most impactful rules are in your agents.md?

    Zeke Sikelianos
    Zeke Sikelianos
    Cloudflare
    @zeke

    Skills™ are great, but you can get a lot of mileage from a few precise edits to your global AGENTS.md file. Here are my Cloudflare pointers that fill the knowledge cutoff gap and give my agents some context about my account setup:

    Image
    45
    Reply
  • When we first demoed Claude Code internally, it got two reactions on Slack. A year after GA, @_catwu and I sat down to talk about what's changed: why I use auto mode instead of plan mode, how routines fix bugs before I see them, why I do most of my coding from my phone now, and Show more

    ClaudeDevs
    ClaudeDevs
    Anthropic
    @ClaudeDevs

    Claude Code's first demo got two Slack reactions. One year after GA, @bcherny and @_catwu look back: verification best practices, why we built auto mode, routines and loops, and what's next. youtube.com/watch?v=Hth_tL…

    2.2K
    Reply
  • introducing loops! a directory of pre-built agent workflows for Cursor, Claude Code, and other coding agents. copy a kickoff. set exit conditions. let the agent loop until the job is actually done. 26 loops live → loops.elorm.xyz

    Image
    Peter Steinberger 🦞
    Peter Steinberger 🦞
    OpenClaw🦞
    @steipete

    Here’s your monthly reminder that you shouldn’t be prompting coding agents anymore. You should be designing loops that prompt your agents.

    1.4K
    Reply
  • The numbers may be a bit extreme here, but unquestionably use-cases have to stratify in the next year or two between model families. We’ll see a split between frontier intelligence for high end tasks and work, and much cheaper models for high volume workloads that can Show more

    Brian Armstrong
    Brian Armstrong
    Coinbase 🛡️
    @brian_armstrong

    Good take My guess is - demand for intelligence is near infinite - but 80% of workloads will be running on 99% cheaper models within 12-18 months - 20% of workloads will still run on latest gen models where IQ maxing is important (scientific breakthroughs, higher level

    314
    Reply

Sunday, June 7, 2026

Saturday, June 6, 2026

  • this is interesting in at least one anecdotal setting I've seen, our Reactor Harness uses 150x fewer tokens than the raw Claude Code equivalent for the task of keeping a model of my local agent usage up to date this 150x improvement is not on a "per request" basis. instead, Show more

    Paul Graham
    Paul Graham
    @paulg

    Curiously enough I did office hours today with a startup that cuts companies' LLM token costs by optimizing requests. They can cut costs by about half, which they split with the customer. So the TAM is a quarter of the model companies' corporate revenue. That's a big TAM!

    65
    Reply

Friday, June 5, 2026

Thursday, June 4, 2026

Wednesday, June 3, 2026

Tuesday, June 2, 2026

  • codex has replaced ChatGPT as my daily driver; not b/c I'm coding all day; it helps me move actual work forward: research, analysis, content, ops, planning, and execution. codex for everything.

    The Information
    The Information
    @theinformation

    Exclusive: After falling behind Anthropic in coding, OpenAI built Codex into one of its fastest-growing products and is now making it central to ChatGPT’s future. Full story: thein.fo/3PT3Huu

    55
    Reply
  • YES-CODE An entire category of software, "no-code", was built under the presumption that code is expensive, difficult, and scarce. Coding agents have forever changed the equation. Code is now cheap, easy, and abundant. I remember @cramforce being asked by an analyst long ago: Show more

    Warp
    Warp
    @warpdotdev

    Our warp[dot]dev site gets 10M visitors/year. We migrated the whole thing from a no-code editor back to code in just 3 weeks. Very few hiccups, and SEO actually improved. Plus, the marketing team is free to use Warp to ship future changes

    799
    Reply
  • Workflows are the biggest upgrade to Claude Code’s capabilities since skills and subagents. I dove deep into it with @sidbid to figure out best practices, examples and more.  I’m particularly excited about the non-technical tasks it enables for Claude Code.

    Thariq
    Thariq
    @trq212

    x.com/i/article/2061…

    4.7K
    Reply
  • If your daughter needs tutoring in algebra, you can probably find someone cheaper than Albert Einstein. Giving every task to GPT5.5 or Opus 4.8 is overkill. Often times you can get the task done just as well, but 10x cheaper and 10x faster with a smaller model.

    Factory
    Factory
    @FactoryAI

    Introducing model routing to Factory. Factory Router picks the right model for every task, automatically. Maintain frontier performance while cutting costs by 25%.

    274
    Reply
  • This is right & is what I've been telling anyone who'll listen (= few). I'd qualify this whole thread even further tho: the step-function bump we saw in coding capability at the end of '25 was a one-off. The models are near the top of the S curve in coding capability.

    John Arnold
    John Arnold
    @johnarnold

    There are several reasons software was likely disrupted first: 1) The work is already digital and performed through text 2) Coding has tight feedback loops that allow for easy testing of whether the AI output worked 11/n

    124
    Reply

Monday, June 1, 2026

  • Some tips to help agents understand your codebase: 1. The source code either needs to be the source of truth, or have something legible as a path to the source. For example, if marketing site content is actually stored in a CMS, you need to either delete the CMS and move that Show more

    Lee Robinson
    Lee Robinson
    SpaceXAI
    @leerob

    Cursor just got a major upgrade! Agents can onboard to your codebase, use a cloud computer to make changes, and send you a video demo of their finished work. The latency of using the remote desktop is smooooth.

    1.1K
    Reply
  • slowly we're all realizing that tools should be called from code, not from within the llm api

    Perplexity
    Perplexity
    @perplexity_ai

    Introducing Search as Code, our new search architecture for AI agents. It writes Python that calls our search stack directly, instead of looping through function calls one at a time. Available in the Perplexity Agent API, and now default in Computer. research.perplexity.ai/articles/rethi…

    Image
    495
    Reply
  • lot of talk today about "vibe coding" / "is vibe coding dead" - its not. The problem is, lots of us, on various timelines, tried to apply vibe coding to engineering our systems in production and realized that *really* doesn't work it's always been a useful thing, I *also* vibe Show more

    dex
    dex
    @dexhorthy

    using AI for coding is a deeply technical engineering craft most people don't approach it as so, and don't get the results we associate with high craft but the ones who do have been sprinting ahead more tokens wont save you, more thinking + skill + llm intuition will have

    150
    Reply