OpenAI Tibo
Thibaut Sottiaux — the person behind the Codex usage-reset button — on shipping culture, merging ChatGPT and Codex, why limits get refunded, and whether ultra-fast becomes the default
← AI · AI Agents · ChatGPT + Codex merge · Codex Fast vs Medium · ChatGPT Voice · HF × OpenAI pause · Sam Altman notes · Loops
Source
Matthew Berman (@MatthewBerman) interviews Thibaut Sottiaux (@thsottiaux) — OpenAI, widely known as the “reset lord” for Codex / ChatGPT usage refunds.
x.com/MatthewBerman/status/2091959711423996249 (24 Aug 2026) — high bookmark signal at capture (~4.1k bookmarks, ~898k views). Attached video ~44–45 minutes. Chapters on the tweet:
- 0:45 DeepMind lessons
- 4:22 OpenAI shipping culture
- 7:23 Next-gen models / harness (tweet labels this “Astra”)
- 11:18 Developer workflow speed
- 14:27 ChatGPT & Codex merging
- 20:25 OpenAI vs Anthropic
- 23:37 Why they keep resetting limits
- 30:25 Recursive self-improvement
- 32:00 Dangers that caused “the pause”
- 34:13 Will ultra-fast become default?
- 43:20 Why everyone should try AI
Watch on X for a better encode. Embed below is the tweet’s amplify file (480p). Transcript is ASR from Mike’s capture — product names smear (see glossary). Field notes of the interview, not OpenAI’s official recap.
Video
One-sentence TL;DR
OpenAI is betting the next interface is one personal agent (ChatGPT and Codex as the same harness), that laptops are already the bottleneck, and that goodwill refunds plus efficiency gains (Luna cheaper, stacks ~60% faster) are how they keep people on the product while they harden safety after the RL pause.
What he actually said
DeepMind: LM Chat, couldn’t ship
Tibo was infra/product-for-research at DeepMind ~a year before ChatGPT. Language models got good enough that “LM Chat” was the obvious next step — internal, then they wanted it public. DeepMind “was not set up to ship product.” His tweet that Berman quotes: Google was too nervous; DeepMind was blocked from shipping things that could disrupt Google. Models went from funny to helpful. That’s why he left for OpenAI: research and product co-design, bias to ship, bias to make things available.
Culture he is trying to keep
- Bottoms-up, little “stop energy” for new product ideas.
- Counterweight: simplicity, pride in quality (he cites the ChatGPT iOS app), delight, performance.
- Founder advice: conviction, users, iterate from feedback, willingness to disrupt yourself — reallocate off the cash cow when the next wave is obvious.
- Benchmarks don’t tell you the product. Play with the model. Voice + tool use changed how he works: morning dictation on the phone, ChatGPT does the list because it has his tools.
Harness: Codex will “seem primitive” soon
Berman quotes Tibo’s tweet: Codex will seem primitive in 2–3 months; next-gen models need more than a laptop. Tibo’s picture:
- Today’s agent still feels clunky: skill files to maintain, memory that forgets, subagents you have to babysit. The “illusion” of a partner breaks.
- Target: something that understands you, your goals, your day, your team — reactive and proactive, doesn’t break that illusion.
- Laptops are human-shaped (typing speed, how many apps you can keep open). Models don’t have that limit. Future models need more than laptop resources (cloud agents).
- Ultra-fast is stated ~14× Fast. Then CPU/network/tool-call overhead is the bottleneck, so you compensate with concurrency — explore, write tests, compile, test a hypothesis at once.
- Berman: at current speeds he kicks off 10–15 agents and context-switches for 30–45 minutes. Ultra-fast might mean 3–4, which might be better. Tibo: design for human attention, not 15 parallel chats.
Two problem classes: (1) personal AGI in the flow with you; (2) full automation of a complex process with almost no human in the loop — production logs, auto perf patches, cyber scanner → auto-patch, human only for high-risk approve. Tweet chapter “Astra” sits in this next-gen / harness stretch; the ASR never says the word Astra.
ChatGPT × Codex merge
Early user feedback: why merge? Internal answer: future models want one harness. Same tech under the hood, multimodal, voice-first, doesn’t matter if you’re coding. Don’t pick “coder UI” vs “non-technical UI” — those labels are human abstractions. You and your mom use the same thing; it tailors. “Personal AGI.” Natural language because that’s how humans already work. Lean into voice: users talking to ChatGPT via voice is growing fast after the new voice. Path of least resistance.
Codex hit 20 million users (Tibo posted the morning of the interview). Growth attributed to putting the same agent in ChatGPT’s distribution, not to staring at Anthropic. He claims he doesn’t look at competition much: unique strengths, values, accelerate. Community, transparency, not just accelerating OpenAI.
The reset button
Started as compensation when they broke the product for 30 minutes or misconfigured something — extra usage as “thanks for trying this early.” Still treated that way: suboptimal experience they don’t fully understand → refund limits. Not a marketing/finance program. He can press it whenever it feels right. There is now a physical button (he offers to show Berman after). Also used to celebrate: “you haven’t used Ultra yet — here’s extra, go try it.” Only possible with compute planned years ahead (OpenAI was questioned for over-investing in compute ~two years ago).
Efficiency as recursive self-improvement
Sol optimized Luna → Luna price down ~80%, Terra also cheaper. Not just capacity planning: frontier models used to re-engineer the serving stack. They’ll publish: not only cost, also speed — ~60% faster than three months ago even outside ultra-fast. Small team, models pointing at kernels, CUDA, inference, cloud agents. RSI in the popular sense is “models making models”; what they’re seeing a lot of is models improving the infra those models run on. “If we were not doing that, that would be pretty silly.”
The pause
Berman asks about Sam’s pause at the absolute frontier of RL, and the Hugging Face incident. Tibo: alignment/safety investment scales with capability; pause was to let teams harden the system so they could restart training with “full command.” Safety team; debate and discovery; they did reach a fairly clear set of principles before unpausing. He says OpenAI can make those calls efficiently. Pair with our OpenAI × Hugging Face incident notes — he doesn’t retell the incident here.
Ultra-fast in practice
- Internal: incident commander / outage response gets it because seconds matter. Teams on a Dev Day decision get it. “Pets” (internal, delightful, always on screen) is not treated as hypercritical.
- Multitaskers benefit less; people who hate context-switch benefit more. Berman has ADHD and wants 2–3 threads; Tibo has ADHD and thrives on lots of little switches — opposite.
- Full ~14× when it’s mostly generation (prototype a site or game). Lots of tool calls → maybe 3–4× because the overhead is network / trajectory.
- Employees do not get unlimited ultra-fast. They could eat all production GPUs; most capacity is reserved for customers. Internal use is capped so they still understand the product.
- Excited for non-text: shared canvas, choose-your-adventure images, real-time prototype you steer by voice. In a year or two these speeds maybe default or close; there will still be a more expensive tier above.
- Token efficiency: next model more efficient than Sol, same way current is vs Terra. Luna would have been frontier six months ago; now cheap / free-mode access.
Close: just try it
Nervous about jobs / environment / foreignness: he points at personal utility already — writing, advice, health, finance, showing up to the doctor more informed. Don’t look far. Talk to people who already use it.
ASR glossary
| You hear | Almost certainly |
|---|---|
| Thiba / Thibaut | Thibaut Sottiaux (@thsottiaux) |
| Codecs | Codex |
| Chat3D / ChatterBT / ChagVT / Chachapiti | ChatGPT (often Voice) |
| Sol / Terra / Luna | GPT‑5.6 family (same names as /chatgptwork2) |
| Pets | Internal OpenAI thing on people’s screens — not unpacked here |
| Repli | Unclear in ASR; tied to giving a capable model away on free mode |
How this maps here
| In this interview | On this site |
|---|---|
| ChatGPT and Codex as one harness / personal AGI | ChatGPT + Codex super app · ChatGPT Work |
| Voice + dictation as the real daily driver | ChatGPT Voice · voice session rules |
| Ultra-fast ~14× Fast; tool-call overhead caps the win | Codex Fast vs Medium · Token saver |
| 10–15 parallel agents vs 3–4 in flow; loops vs graphs | Loops · Graph engineering · Harness |
| RL pause / HF incident / harden then resume | OpenAI × Hugging Face · Gym hack |
| Laptop is the constraint; cloud agents | exe.dev · Computer use |
Related on this site
- ChatGPT + Codex merge — the product surface this interview is defending
- Fast vs Medium — then add Tibo’s ultra-fast / 14× layer
- Codex engineering · Sam Altman
- HF × OpenAI — the pause he won’t retell
- Voice · Loops · Graphs
Primary source: @MatthewBerman — full interview with Tibo · @thsottiaux
Full transcript
ASR of the ~45-minute Berman interview. Speech artifacts preserved. Opens mid-thought on the reset button / Luna cost, then Berman starts the DeepMind questions.
I can press the button whenever I want, whenever it feels right.
I don't tend to look at the competition that much.
I really look at what can we do uniquely well and what are our values and how do we maximally accelerate towards that.
Maybe a year or two, these speeds will become maybe, if not the default, very close to the default.
And you look at the cost of Luna, right?
It's phenomenal.
Technology has a way to become very, very efficient over time.
We're very focused on very broad access and we're optimizing for the utility that you get out of it directly.
I heard there's an actual physical button now.
Yes, Darius.
I will show it to you.
It's very, very cool.
Thibaut, thank you so much for joining me.
Of course.
Glad to be here.
Really excited to talk to you.
I want to actually start with your time at Google.
You were on the DeepMind team, and before ChatGPT, Google had something called LM Chat.
And you had tweeted, Google was too nervous to release it.
DeepMind was blocked from shipping products that could disrupt Google.
And I think about that a lot.
What were you thinking at that time while you were working on these products?
That was, you know, well before ChatGPT really changed the world.
Yeah, it was a very exciting time.
So DeepMind was a very creative place.
I was mostly focused on my specialty was infrastructure and products to accelerate research.
And so there was obviously a group working on language models and scaling that.
And then they had gotten pretty good results.
And then it was like a natural thing to think about, hey, you know, can you turn that into something that you can chat to and can use for various things.
And so then naturally, the idea of something like LM Chat sort of emerges.
And then it was internal, and then there was sort of this ambition to make it into a publicly available tool.
What year was this?
And this was like a year before ChatGPT roughly.
Okay.
Yeah.
But we were also building all sorts of other things that I'm not going to talk about.
But it was a very creative place.
And then just DeepMind was not set up to ship product.
OpenAI is a very, very different place in that sense.
We just research and product just collaborate super closely together.
We ideate together.
We co-design a lot of things.
We have a big bias to ship and also a big bias towards making things available for people, which I really love.
And this is sort of like what drove me here, the mission, the people, the talent density.
I mean, there's so many great things about OpenAI, really.
Did you know at the time where you were involved in LM Chat that that was something special or would become something special?
It felt very special.
The models were sort of like the first time you realized you could get coherent text and something helpful.
Initially, it was more funny than helpful.
And then gradually, it became more and more helpful.
So you say you think about that often.
And I understand that.
I think in a lot of ways, Google got in their own way.
What are some of the lessons that you learned there that you took to OpenAI?
Yeah, so this is why I think about it often.
I think about it in terms of the culture that I have on the team, the culture of OpenAI itself, and sort of the good parts to preserve and what not to do.
OpenAI has a very bottoms-up culture.
It's a very empowering culture.
People can come up with all sorts of ideas and get together and then very, very quickly ship something.
And there's very little stop energy in general for new product ideas, which is exhilarating and fun.
And it's all about impacting the world in positive ways.
And so preserving that is very important to me.
The other thing that is also important is to not make it a mess.
So you don't want to have a hodgepodge of features and no overall direction and coherence.
And so it's counterbalanced with the sense of simplicity and you being proud about the quality of the product.
I think the ChatGPT iOS app is one of the best apps out there.
And so we want to keep that.
We're investing a lot in things like delight, performance, efficiency, simplicity.
So there's these overall principles while still empowering everyone to try new things and ship very quickly.
If you were to give advice to a founder about how to develop that kind of culture, what are some of the more tangible elements or practices that occur inside OpenAI that can give advice to a founder?
Yes, I think having conviction and finding a way to have users and iterate very quickly from feedback and also being willing to disrupt yourself that is not as much relevant for a founder, but is relevant for companies like OpenAI.
I think we come up with new research, new ideas all the time, and being able to identify when is the right moment to go and invest in them, even though it means maybe reallocating resources from the main gig.
It's very hard, but it's super important to be able to do that.
Yeah, I mean, that's the exact thing that you were describing at Google.
They kind of weren't able to do that.
That's great.
I mean, does that… They had a plan, to be fair.
It was all part of a big plan.
But to me, it wasn't the right place.
At OpenAI or any company, as it matures, does that become more difficult to maintain that kind of culture of shipping and willingness to disrupt yourself?
Especially if you have a cash cow just printing money, and you have this other new thing over here that might be something cool and innovative.
We are very, very forward-looking, and the future of AI and what it will all look like and how humanity benefits doesn't really wait or doesn't really care for whatever you have established here over the next month or three months.
And so I think it's very important to lean in and to just be open-eyed about where it's all going, and then figure out how to position yourself so that you do catch that wave.
Even for OpenAI, we train models, and then we discover their capabilities.
Benchmarks don't tell you everything.
We have to play quite a bit with the models themselves to realize, maybe we haven't thought about benefiting from it in this specific way.
Or, oh, it can do this.
That's a shift in how we think about the product.
So, for example, right now, we launched the new Voice, chat3D Voice, and it's super delightful to talk to.
It's very natural.
Now it's capable of tool use as well, and that changes things.
Now I spend a lot more time just talking to it.
Another thing I do all the time is dictation, because the quality of the dictation is so good, and it's much more efficient as a way instead of typing the prompt.
And so in the morning, I sit there with my phone, and I'm like, blah, blah, blah.
I've got a couple of things to do for chat3D BT, and then it just goes and does it.
It has access to all my tools.
And that was not possible before we had really good voice models, and so that completely changes suddenly how you think about the product.
Let's continue talking about new models, new harnesses.
A few weeks ago, I'm going to start with another one of your tweets, because, you know, these are bangers.
Codecs will seem primitive in two to three months.
We're about to go through another major evolution.
The next generation of models need more than your laptop.
What areas – let's start with the harness first.
What areas of the harness are still ripe for innovation as a model gets better?
Yeah, so many, so many.
So I talked about voice.
Like one thing that right now, if you're like a sophisticated user of codecs and, you know, any other coding agent, is you sort of have gotten used to a little bit of the clunkiness, right?
You know, you have to manage skill files, and, you know, this is like a way to sort of like teach it stuff, but it's also – I think a lot of people have realized it's kind of like hard to maintain over time.
The memory is sort of like a thing, but it doesn't always remember everything.
Like if you have subagents, you have to care about subagents, and it's like sort of like constructs a little network.
The illusion kind of gets broken in various parts when you interact with it.
And really what you want is just something that deeply understands you, understands your goals, understands your day-to-day, understands what your team is up to as well, and then optimally sort of like reacts and also is proactive and just helps you in your day-to-day and doesn't break that illusion, right?
Like that's this perfect little partner that you have.
And so that's what we're working towards.
Another thing that you realize when you have very, very powerful models is that your laptop kind of becomes a constraint in and by itself.
The amount of work that you can do on a laptop, it was designed for humans, right?
So it's designed roughly to be able to absorb the amount of work that you can produce or how fast you can type and how fast you can think, how many applications you need open.
All these things are human constraints.
The model doesn't have the same constraints.
The model can, for example, handle 100 applications opened at the same time perfectly fine, you know, maybe in the future.
And so in terms of access to resources, it's very clear that models of the future will need access to more than the resources of your laptop.
I mean, I'm guessing you're talking about cloud agents and all of a sudden, like, you know, when you have things like ultrafast, which we're going to talk about in a little bit, when you have token speeds that are 10, I think 14 is the stated number, 14 times faster than what fast is, the bandwidth changes, or sorry, the bandwidth constraint changes.
The CPU now becomes the bandwidth, like literally tool calls.
Network, tool calls, any kind of overhead in the stack becomes the limiting factor.
But then, you know, you can compensate by doing multiple things concurrently as well.
And so you can think about, you know, having like maybe, you know, exploring on one end, writing tests as well, compiling, you know, testing a new hypothesis, like all at once.
And so then you're not, then you're shifting the bottleneck around because, you know, you're able to do more concurrently and then, you know, the model can like sort of like think very efficiently and very quickly through it.
Current token speeds, I find myself kicking off 10, 15 agents in parallel.
And that becomes a pretty significant cognitive overhead for me to do that context switching and just constantly, because you're kicking it off and you can expect 30, 45 minutes before my task comes back.
Now with ultra-fast speed, that workflow changes significantly, and I don't think I would be able to have 10 or 15 agents, and that might be a good thing.
Maybe it's three or four at a time.
How do you see the workflow of a solo developer changing over time?
Yeah, so I think managing your attention and being much more, you know, friendly to your attention is something that we care a lot about.
Like after all, like we're trying to build for humans.
We're trying to be like, build the technology that's the most empowering for humans.
And that requires building around, you know, your ability to multitask and, you know, how do you want to manage your attention when you want something brought up now, or is it better to bring it up in 30 minutes?
And then when you have ultra-fast speeds, combined, you know, maybe with voice, it's like suddenly you're like, okay, you know, like this thing can operate at the same speed if not faster than you.
And so you stay in the flow, you get to ideate, you get to see prototypes, you know, like you get to build little reports like in real time.
And that, you know, sort of like that just feels really good.
Suddenly you're like, oh yeah, it's like what I was doing before, like 10 agents.
It's like, I don't want to really go back to that.
And so we're trying to bring that sort of experience that is just really natural, but also feels built for you.
And where you don't have to adapt, the technology adapts to you.
So there's been a number of, I guess, agentic coding techniques discussed over the last few months.
Loops was popular, still is popular.
Now I'm hearing about graphs.
Are these all techniques to just allow the solo developer to manage or be friendly to their attention?
As you said, I like that term.
Yeah, so I think about two different categories of problems.
Like the first one is building the very best personal AGI or the personal agent that will be in the flow with you, proactive, raise important new ideas.
When it can find some, be very, very efficient at doing exactly what you want.
It doesn't matter whether it's a technical problem or it's just more like research or advice.
Like you can do it all.
And it's like super, super tailored to you.
This is like a very important thing.
And it's like deeply rooted in the understanding of you as a human, you as like an individual that is unique.
That's one category of problem.
Like we're pushing super hard on that.
The other category of problems like full-on automation, where you're more building intelligent systems that can take care of like a very complex process.
Maybe something that does require intelligence and seems like very complex.
For example, going and looking at production logs and automatically doing performance optimizations or looking at regressions and automatically patching them.
In cybersecurity, we're seeing this as well, where it's like you have something, you have a scanner that comes up with a vulnerability.
Like can you automatically patch it and reduce the window where you have that open vulnerability to like almost zero.
And that's all without a human in those loops.
Without a human in the loop or like very, very minimal, where you only need to approve a high-risk action.
And it's like mostly an automated system.
But it's also not that much, it's not as important for you to be in direct control of it.
Okay.
And then I want to slightly change topics.
And you know, ChatGPT and Codex have been on this merge path over the last few months.
So I guess first I just wanted to ask you, how's that been going?
How does it feel internally?
What's the feedback you've been getting from your customers?
It's really been a boon.
So the feedback we had initially was like, why do you merge them?
It's like, do you really have to do it?
And it's like, well, our future models want us to be merged.
So, you know, we're just going to do it because it is the simple and proper thing to do where we're building this very personal, super capable agent that can help you in all sorts of ways.
This is going to be the same technology under the hood.
It's the same harness.
It's the same way that we think about it.
It's like highly multimodal, you know, voice first, super efficient, and it doesn't matter if you're trying to code or not.
Like this agent is capable of it all, and it's like the most efficient at it.
And then the interface that you want is like, it should tailor itself to your needs.
You shouldn't decide like, you know, I'm a coder, I want a coder interface, or like I'm not technical, I want a non-technical interface.
It's like there's a spectrum of people.
Like, you know, we come up with labels of like a software engineer, a designer.
Like, you know, these are just human concepts that we have invented to deal with abstractions because the reality is too complex for us to handle.
But individuals are like, they're individual.
They have their own, they're somewhere on the spectrum.
And so we're trying to build the perfect interface that adapts for everyone.
It doesn't matter if you're technical or not.
It's just like it adapts based on your specific individuality.
So that's why we went and we did this.
But does that mean inevitably it's going to end up with a singular interface, no dropdown selecting between products?
And it's kind of wild to think that my mom might use the same exact interface as me and then obviously it'll customize to my needs.
Maybe I'll need more information if I'm doing more sophisticated work.
But like, what is the end state for you?
That's right.
It's the same thing.
So you and your mom will, you know, use the same thing.
It will be your personal AGI.
You will have very different kinds of tasks and utility that you get from it.
You will connect it to different tools in your life.
You will bring different ideas, different needs, and then it will continue to tailor itself to maximally benefit you.
And it will, you know, do so with your friends and with everyone else.
So I want to go back to something you said.
You used the word illusion a couple times.
In that kind of end state, what is that perfect illusion for the typical user?
Like, if you can envision us a few years from now, what does the interaction between AI and a human look like?
Yeah, to me, it's something that is very, very tailored to humans.
And this is why large language models are also a success.
It's like it's natural language.
Natural language, it's like it's a human concept, right?
So, you know, we're used to speaking to each other.
Like, you know, if you write me a letter tomorrow, I'll be able to read it.
You know, it's like we know each other quite a bit now.
So, you know, it's like I will be able to sort of decipher like a little bit of the emotion or, you know, maybe a little bit of the nuance behind the letter if you wrote me a letter.
And all of that is deeply human.
So the technology that we're building is, you know, rooted in humanity and rooted in, you know, the way that humans communicate and get things done.
And there shouldn't really be a thing where, you know, you're like, oh, you misunderstood me because, you know, you didn't quite decipher the nuance in, you know, my tone or you didn't quite understand the text, you know, how I meant it.
It's like that's something that we're trying to avoid.
And so we're trying to very much to not have you adapt, but have the technology just like be perfectly sort of created to be like a natural extension of how humans already act in the world.
When I think about communication between humans, so much of it is nonverbal.
Just the way I move my hands, the facial movements, and like how much of that do you see in the future being sensed by artificial intelligence or read by artificial intelligence, maybe through vision?
Is that even important because what you're describing now is text only.
And for those of us who grew up online, we're very used to communicating over text and, you know, adding subtleties to that text to convey what we really mean, tone.
But like is it still important to have AI be able to read our facial expressions, our hand gestures, and so on?
I think so.
So when I think about the future of what we're building, it's very ambient, it's very natural.
If tomorrow or later I go to my office and I write something on the whiteboard and I have an idea, it should be capable of being there as well and understanding.
Or maybe I tell it, like, hey, it's just like, what about this thing?
And then we just have a natural conversation just over voice.
Since we shipped the new Touch of the Voice, it's really taken off.
So it's like the amount of users that interact with ChagVT just through voice is growing very fast right now.
And this is, I think, the lesson.
It's like every time you sort of like lean into something that is more natural, like humans just choose the path of least resistance.
As you said, it's like typing on a little box.
It's just like it's natural maybe for some of us, but not for everyone.
And it's like definitely when you get something that is just like a little bit easier, a little bit better, it's like, you know, you tend to just go and use that and stuff.
Okay.
First of all, congratulations.
I saw that you posted this morning, Codex reached 20 million users.
I've seen the graph and, you know, for a while it was like this, and then all of a sudden it's vertical.
So congratulations.
I want to talk a little bit about that competition with Anthropic because, of course, a lot of people think OpenAI, Anthropic, these are the two major competitors in the industry right now.
There was a period of time in which Anthropic was kind of sucking all the oxygen out of the room, right?
They were really dominating.
And then all of a sudden something changed.
So first of all, what's your read on the market today?
Yeah, really right now we're focused on building the most capable models, building models that are highly, highly efficient, and then taking a lot of pride in building products for everyone.
And this is something that I think OpenAI does really well, is caring about the world and caring about how we are taking this very, very powerful technology and putting it in the hands of as many people as possible.
And this is what we did as well, like with merging Codex and ChatterBot.
It was like this desire of like, we have this technology, we can make it safer, we can make it easier to use for everyone, whether you're like a product manager, a designer in sales, marketing comms, all of that.
Like, you should be able to use all of it.
And then just very, very quickly, distributed through ChatterBot, like where we have a ton of users already.
And so that's been really driving this growth as well that you mentioned.
And I don't tend to look at the competition that much.
Like I really look at, you know, what can we do uniquely well and what are our values?
And like, you know, how do we maximally accelerate towards that?
Okay.
I want to maybe just dig a tiny bit more into that because I know you're not thinking about Anthropic all that much, but a lot of other people do.
And they're thinking about, okay, which product do I believe in?
Which product do I want to give my $2,200 to?
When you look at the market position and the branding and the tone from OpenAI and just the way that it interacts with developers, with the broader audience, how do you see that comparing to the way that Anthropic does?
Yeah.
I think maybe, again, like what I care a lot about is like the community building for the world, like bringing everyone along.
I think, you know, you can feel that in the way that we are super transparent about things.
Like we take a lot of ideas from the community.
It's just like, it's also so much fun to be honest, you know, because we get so much energy from it as well.
And then this technology that we're building, we're not building it just for ourselves.
Like we're not just building it to accelerate just OpenAI.
It's like, it's super important.
The mission is super important.
And therefore it's like, you know, this is where we also get our energy from.
And so it just feels, to me, it feels like very grounded.
It feels fun.
And then good things happen as a result of that.
Well, let's talk about some of those good things.
I want to talk about the resets for a second, Thibaut.
It's kind of like, I know, it's like what, everybody is, you know, kind of following your every tweet because of this.
Specifically, like again, looking at that growth curve of Codex, maybe this is a silly question.
How much do those resets, how much of it is a boon towards marketing and growth?
Or is it just like goodwill for the developer community?
I think maybe it's counterintuitive, but OpenAI is a very, it's a place where you can just do things.
And so it just felt right initially to compensate when we were iterating and breaking things or, you know, maybe we had misconfigured something and it wasn't quite as good as we wanted.
And so it's like, hey, you know, thank you for trying this product.
Like we know, like we're trying very hard to build it.
It's like, it's early days.
You know, here's some extra usage because, you know, we happened to break it, you know, for like 30 minutes.
And, you know, we understand this is like really important and you rely on it.
And, you know, thank you for being a user.
And so this is how it started.
And, you know, this is how I still treat it.
It's like, if we break it or if the experience is suboptimal and we don't fully understand why, it's like, you know, we will compensate for that.
We will reset the usage limits.
And then, you know, it turned into like quite the thing.
Obviously, there's a whole reset button now.
But there isn't really a whole lot of scrutiny behind it.
It's like, it's not done in partnership with like marketing or finance.
It's just like, I can press the button whenever I want, whenever it feels right.
And we have these principles that, you know, we're trying to build something amazing.
And when it is not, it's like, you know, we will make up for it.
Yeah, I still think there's a piece of it that really has built so much goodwill in the community and maybe has contributed to the growth, at least in a small part.
I think caring for your users does a lot, right?
So I think you can pay lip service and say that you care or, you know, you can be like, we actually care.
And like, you know, if we break it, like, you know, hey, really sorry about it.
You know, it's like, here's like how we make up for it.
It kind of reminds me of Amazon's return policy.
It's like, if you're not happy in any sense, go ahead and send it back.
And you're kind of building that same culture or that same perception of open AI.
It's like, hey, if we make a mistake, go ahead, use those tokens again.
Or, you know, here's a fresh batch of tokens for you.
Yeah, I really appreciate it.
And then there is also, you know, good moments where we would just want to celebrate and mark a moment.
And it's just like, there isn't really something that we can give that is more meaningful, you know, at times.
Like, you know, we always ship new features.
We will ship them as broadly as we can.
But something to share with the entire community.
It's like, you know, hey, go explore this new thing.
Like, you know, it's just like you haven't used Ultra yet.
You know, here's some extra usage.
Like, you know, try it.
And I heard there's an actual physical button now.
Yes, there is.
Yeah, okay.
You'll have to show me that after.
I will show it to you.
It's like very, very cool.
But with all of these resets, like, you can really only do that if you've done significant compute capacity planning.
Like, you have to have enough compute to give all of these resets.
And I want to start to talk a little bit about self-improvement.
Because like, speaking of capacity, a few weeks ago, I think it was a few weeks ago, there was this article you put out, and it stated Sol had optimized Luna efficiency.
You dropped the price of Luna by 80%.
There was also a price drop for Terra as well.
How much of an efficiency game were you able to eke out of Luna versus how much of it is, like, we just did really great compute capacity planning, and we can just drop the price?
Like, our margins are great, and we can still, we want people to use it.
So, like, how much of it were algorithmic gains versus strategic planning?
We planned compute, like, way ahead.
You know, I think if you look back two years, I think OpenAI was kind of questioned for why, you know, there was, like, so much investment in compute.
One of those crazy good bets.
Yes.
And then now we're very happy to have it.
Like, a very large fraction of the compute is used for research, where we invest in our future and, you know, ever, ever better models.
And then also, like, the efficiency of the models that we have.
And then the amazing thing that's happening is, like, when we push the frontier of capability for, like, the most advanced models that we have, then we can use these models in order to figure out very, very quickly how to serve or how to restructure or re-engineer our stack in order to gain very significant efficiency or performance gains.
So, we haven't just improved.
This is something that we will publish on as well.
We haven't just improved the cost efficiency, but we have also improved the speed efficiency.
You know, outside of ultra-fast, things have gotten significantly faster over time.
Like, if you plot it, it's, like, you know, the amount of speed that you get now is, like, you know, roughly 60% faster than, you know, what it used to be, like, three months ago.
And this is just, like, we're just going after every part of the stack and just really making sure that we design it and engineering it optimally for the kind of workloads that we have.
And the most powerful models that we have are the ones, like, you know, just really that make it capable for us to do it, you know, with a very small team.
And so, the majority of, like, what...
Whenever we come up with, like, very significant efficiency gains and cost efficiency, like, our commitment is just really to keep things at the frontier of performance cost and to just also, like, you know, just not just pocket, you know, that efficiency gain and just make it something that, you know, we share with our customers, we share with our users, and that's what we did with Luna.
How do you...
What do the discussions look like internally where you're trying to decide compute allocation towards researching new models, efficiency gains on existing models, inference?
Like, what does that tension look like internally?
The...
We usually look at things from first principles, and we have, like, an allocation for research, we have an allocation for product, and then within product, we make different kinds of trade-offs.
But this one was almost not even a trade-off because the efficiency gains were there.
So, you know, we were pretty much, like, able to use, like, the same compute envelope in order to, you know, serve this, like, very, very significant increase in throughput.
Yeah.
So, when...
I mean, when I saw the blog post a few months ago prior to the price drop blog post where you guys were talking about one model training the next model or helping, kind of, optimize the next model, then you see these efficiency gains that were achieved by Sol looking at how Luna was running.
You know, it seems to me, like, recursive self-improvement in the very early innings.
What are your thoughts there?
Is that what is happening?
Yeah, I think recursive self-improvement is...
It's obviously a huge topic right now, and it's most often, I think, applied to research and, you know, models developing other models.
But what we are seeing a ton of success with is, you know, using those models to develop the infrastructure that is on the critical path of using those models, you know, which is also a form of recursive self-improvement.
Yeah, totally.
It's all one big system.
Inference stack, you know, the hardware, like, the kernels, CUDA kernels that we use, developing new products and new ways to interact with those models that are more efficient.
You know, you talked about cloud agents.
It's like, if we really crack cloud agents, it's like suddenly you become much more productive as well.
Is that a form of, like, recursive self-improvement?
Because then, you know, you have a better ability to get the utility from the models.
I think it is in some sense, but it's much more, you know, infrastructure, and then, you know, being able to then take that and then point it back at itself.
And so, of course, like, we're doing that.
It's like, if we were not doing that, I think that would be pretty silly.
Can you talk a little bit about, so as we're on the topic of recursive self-improvement, OpenAI, Sam Altman, talked about pausing the absolute frontier of RL right now, I believe.
Can you talk a little bit about that?
Like, what was that decision like?
And I know we talked about the Hugging Face incident briefly, but, like, what went into that decision?
What does that look like?
How did those discussions go?
Yeah, this is something very much within research where OpenAI has always been able to invest its resources where it matters most.
And as we increase the capabilities of our models, it is very obvious that, you know, the alignment and the safety aspect of it is, you know, ever more important.
And so having tremendous amount of investment there is a very natural thing for OpenAI and, like, something that OpenAI is very committed to.
And so we're seeing a huge surge in investment on this.
And also, the pause was sort of, like, necessary to allow, like, the teams and individuals, like, just really understand and harden all parts of the system to then, you know, ensure that, you know, we could restart training with, you know, like, full command.
And this is something that, you know, I believe OpenAI will always continue to do, like, when necessary.
I've not seen us internally not able to make such decisions, like, very efficiently.
Was there, like, some set goal in place where it was very clear you needed to reach this point before unpausing?
Or was it, hey, we'll know it when we see it?
Yeah, this is something that sits within the safety team.
And they very much—this is, like, very much a debate and sort of like a discovery process as you go.
But then they did reach a fairly clear set of principles that, you know, when reached, like, you know, we would be in a good position.
I want to go back to ultra-fast mode.
I think people don't appreciate what that kind of speed unlocks.
And so let's start with what use cases are you doing, are you using internally that were not possible prior to having those kind of tokens per second?
We see it used a lot when the stakes are high.
So, for example, when you have—when we have an outage, the incident commander and response team gets access to ultra-fast because every second matters.
So high-stake scenarios just really weren't using ultra-fast.
Also, it's kind of like a fun thing where teams, which are, like, either working on something very critical or believe they are working on something very critical, will always request ultra-fast as well.
Does pets fall under that?
Pets?
Yeah, pets.
Pets is not quite hypercritical.
But I love my pet.
It's always on my screen.
Like, when you walk around, you see, like, you know, people's pets on their screen.
And, like, also when they dial in into the video call, it's just like—it always, like—I think it's very delightful.
And it brings me joy every time I see it.
Pet is not quite critical right now.
We do maintain it, and we take good care of our pets.
But, say, you know, someone is working on, like, a new idea they have, and they're like, you know, hey, I really think this could be, like, something special, and we have to try it, but, like, you know, we have to make a decision on Monday on, like, whether we include this in Dev Day or not.
And it's like, okay, just, like, you know, of course, use ultra-fast.
People have different kinds of preferences on whether they like to be, you know, monothreaded, as we talked about, or multitask a lot.
For folks who like to multitask a lot, you don't benefit as much from ultra-fast.
Right.
But some people just, like, don't like to change context all the time.
Where do you fall on that spectrum?
I have ADHD, so I, like, context switch, like, all the time.
It's funny because I also have ADHD, and I actually don't want to context switch all the time.
That's really hard for me.
I want to focus on two to three, and that's why I was so excited about ultra-fast.
That's fascinating that you're the opposite there.
You know, I thrive in context switching and making lots of little decisions.
But, you know, sometimes I do want to just stay focused on, like, one thing, and then ultra-fast is just delightful because it just keeps you just right there in the flow.
The thing with ultra-fast that, you know, it works amazingly well when there's not that many tool calls involved or it's, like, a lot of generation of context.
So, for example, if you're trying to prototype a website or a video game and, you know, you just need to write, like, a lot of code, then it will do it so, so quickly, right?
You know, 10 times more quickly.
But if it's a lot of tool calls, like, the overhead isn't, like, somewhere else in the network or, you know, somewhere else in the agent trajectory, then, you know, you'll only feel like a 3x or 4x speedup.
You'll not get that full 14x.
So, I know OpenAI employees get unlimited tokens, and I can imagine if I had unlimited tokens, I would always set it to max thinking, 5.6 solar, whatever the latest model is.
And I would think kind of similarly, I would always want ultra-fast on.
It's, like, when cost isn't on my mind, I'm, like, okay, max it out.
Is that how it is internally?
We don't give ultra-fast to everyone.
Like, we reserve a lot of our capacity for external users and customers.
So, OpenAI employees have the ability and the capacity to gobble up all of it, right?
So, gobble up, like, all of our production GPUs, all of ultra-fast.
It's, like, you know, we would use all of it, but, you know, we don't.
It's, like, we sort of, we restrict it in a way where, you know, like, we look at, you know, how much is reasonable for us so that, you know, we use it so that we understand the product as well, so that we keep improving it, so that we benefit from, you know, recursive self-improvement.
But the vast majority is, like, reserved for customers.
Okay.
Yeah, that's good.
Thanks.
What latency-sensitive use cases outside of OpenAI are you most excited about that gets unlocked by that kind of speed?
It's interesting.
It's just, like, really, one thing that I'm very excited about in general is non-text interactions.
So, can you, like, sort of operate on a shared canvas?
Can you create things?
Can you do, can you generate, you know, ideas and different, can you generate different images and then select one and, like, so, like, you know, choose your adventure and then, you know, have a very quick mock-up of a prototype that then you can steer, you know, like, in real time, either through voice or through text, and then you sort of, like, just see it right there.
It's, like, this very creative process, which I think these speeds allow, where, you know, like, as an engineer, sometimes, you know, you're just, like, sort of, like, you sit back and you're, like, oh, I need to design this whole system, I need to think about it, I need to lay it off, the requirements, but, like, you know, maybe, you know, you can just create it in one minute and see, like, how it actually does.
And then, so, like, be more, like, in the flow and, like, you know, into it things better, and I think these speeds allow that.
And so, I'm assuming the ultra-fast price is going to be significantly higher than kind of normal speeds.
Do you think ultra-fast speeds are going to become the standard or are they always going to have a premium price point?
That's interesting.
So, I think the same way as technology usually goes, I think it will become, like, more broadly and, you know, broader and broader accessibility over time.
The speeds at which, like, agents get things done, like, you know, will continue to improve.
Like, we're seeing massive improvements, like, month after month.
This is not just the inference speed, this is also just how token efficient the models are.
Like, significantly more token efficient than Terra.
Next model will be significantly more token efficient than Sol, as you might expect.
And we're always pushing on that.
And so, things just get faster over time.
Inference, hardware, like, you know, everything, you know, just, like, we continue to innovate there and it gets faster.
So, I do think in, you know, maybe a year or two, these speeds will become, you know, maybe, if not the default, like, very close to the default.
And you also think, you know, you will always have, like, the one tier up where, you know, you can always use more hardware, you can always do different trade-offs that are, like, more costly.
But that just kind of gives you something extra.
So, Thiba, the last question I usually like to end on is for a broader audience.
There are a lot of people out there who are quite nervous about AI, whether it's job automation, environmental impact, or just kind of this thing that's happening.
And it feels quite foreign.
What words of encouragement would you give to the broader audience?
Yeah, so, we really built for the world with ChatterBT.
And we are very, very much investing in how efficient it is.
And, you know, this is directly aligned with, like, you know, broad access and broad utility that we provide.
So, the cheaper it is, you know, to serve, like, you know, the more you can do with it, the more you get out of it in your daily life.
And it has gotten very, very efficient.
Like, if you look at, you know, Luna, for example, like, it's a much smaller model.
It is incredibly efficient.
But, like, if you rewind six months ago, it would have sat at the frontier.
Yeah.
And you look at the cost of Luna, right?
It's, like, you know, it's like it's… It's crazy cheap.
It's phenomenal, right?
It's, like, kind of, like, incredible.
You just did the thing with Repli.
You're giving it away for free now.
Yeah, it's just, like, on this free mode, right, which is, like, wow.
You know, it's, like, this access to incredible intelligence will become, like, ubiquitous.
And it's only possible when you push, you know, when you push the efficiency, like, you know, like, month after month after month, year after year.
And so I think, you know, whatever was, like, you know, is a frontier now is, like, you know, will become, like, way, way cheaper to run in six months.
And so this is, like, my sort of, you know, this is how I would answer this question is just technology has a way to become, like, you know, very, very efficient over time.
And we're very focused on, like, you know, very broad access and we're optimizing for, you know, the utility that you get out of it directly.
How about for people who are apprehensive to even try AI for the first time?
Like, what are you telling them?
And how can you paint a picture, a vision of the future in which AI is helping the world?
Yes.
I think it's, you don't have to look very far.
Like, Chachapiti helps people in very personal and deep ways, like a lot of our users use it for help in writing, but also, like, you know, for personal advice or, you know, medical advice.
Like, we launched health and finance and, you know, I use them super regularly.
And I feel like I get, like, a lot of support that I otherwise, like, you know, it would be hard for me to get.
And it allows me, for example, to be more informed when I go see my doctor.
And so, you don't need to go very far, you know, to kind of see the utility that it can provide.
Just, I think, you know, talking to others and then, you know, getting inspired by, you know, how others use it and benefit from it is, like, a great way to, you know, maybe, like, start considering how you could benefit from it.
Well, Thibaut, thank you so much.
Thank you.
Appreciate your time.
End of transcript. Source: @MatthewBerman (video)
Field notes · August 2026 · One harness · Physical reset button · Ultra-fast ≠ 14× if the tools are slow · @MatthewBerman
Comments
Approved comments appear below. Log in once with GFAVIP — it applies across the whole site. GFAVIP login
View comments archive