A few months ago, I wrote about why I enjoy running AI at around 20 tokens per second.
At that speed, I can actually watch what’s happening. I can see where the model is going, catch it when it starts drifting, interrupt it, add something it is missing, or just sit there and think while it works.
I’ve realized since then that the speed wasn’t really the point.
The point was that I was still there.
And that led me to another question: if I’m going to stay in the loop anyway, how intelligent does the model actually need to be?
Every few months we get another frontier model. Bigger, more capable, more expensive, better at benchmarks. Then come the charts, reasoning scores, context windows and comparisons.
I look at all of this and increasingly think: cool. But what am I supposed to do with all that intelligence?
Sometimes I absolutely want the smartest model
There are cases where I get it.
Coding is the obvious one. If you’re working across millions of lines of code, trying to understand dependencies across a huge codebase, debugging something complex, or asking an agent to operate with very little supervision, I can understand wanting every bit of intelligence you can get.
The same goes for difficult math, unfamiliar problems, or tasks where neither you nor the information you’ve provided gives the model much to work with.
Sometimes you need a bigger brain.
But that’s not most of what I do with AI.
I use it for research, writing, learning design, exploring ideas and trying to make sense of problems. For that, give me a good 9B to 35B model, curated knowledge, clear instructions and my own brain in the loop.
I’m fine.
In fact, I often prefer it.
The real distinction is collaboration versus delegation
I think I’ve been framing this as a model-size question when the more important distinction is something else.
Am I collaborating with the model, or delegating the work to it?
If I tell an AI, “Research this topic and write me an article”, I’ve delegated a huge amount of thinking. It has to decide what matters, find the angle, determine what to research, evaluate it, build the argument and produce the artifact.
In that workflow, of course I want as much model intelligence as I can afford. The model is carrying most of the cognitive load.
But this article didn’t start that way.
I came into the conversation with an opinion. I’d been watching models get bigger and more expensive, running smaller ones in my lab, and noticing that I often didn’t feel the differences the benchmarks told me I should feel.
I already suspected that context and curation mattered more to my work than another jump in general intelligence. I also knew where the argument might fall apart. Coding was the obvious problem, so I brought that up myself.
Then I basically said:
“This is what I think. Prove me wrong.”
That’s collaboration.
I’m not asking the model to decide what I think. I’m bringing my thinking to the conversation and asking it to push back, find the holes, connect things I haven’t connected yet, and help me take the idea further.
The more thinking I outsource, the more intelligence I need from the model. The more thinking I retain, the more I need a good system around it.
We don’t give AI enough of ourselves
A lot of prompting still amounts to giving AI a topic and expecting it to supply the point of view too.
“Write about X.”
“Research Y.”
“Make a presentation about Z.”
But what do you think about it? What have you noticed? What have you already read? Where do you disagree? What are you unsure about?
That’s context too, and I think it’s some of the most useful context we can give a model.
Without it, the model has to invent the angle and manufacture an intellectual fingerprint for the work. Then we’re surprised when the result feels generic.
I don’t want AI to decide what I think.
I want to bring my thinking and have AI help me push it further.
A bigger brain can still be badly fed
I’ve started thinking about this as nutrition.
Imagine two researchers.
One is brilliant and has read an absurd amount of material. I give that person a vague description of my problem and tell them to figure it out.
The other is simply good at their job, but I give them the 12 papers I think actually matter. I explain what I’ve already researched, what decisions have been made, what I’m trying to understand and where I think the interesting problem is.
I wouldn’t automatically bet on the first researcher.
That’s basically how I think about models now. We obsess over the brain. I’m increasingly interested in what we feed it.
And feeding it well is really a problem of curation.
Even “context is king” doesn’t quite describe it, because you can have too much context. Throwing thousands of documents into a huge context window isn’t good nutrition. You just gave the model a bigger pile to digest.
I’d rather give it the right information.
So maybe curation is king.
That leads me to a different question when I’m running models locally. Instead of asking “What’s the smartest model I can fit on this machine?” I’m starting to ask:
What’s the least model I need to do this well?
I don’t feel your 1%
Spend enough time around local AI communities and you’ll find endless comparisons between Q4 and Q8, benchmarks moving a point here or there, and charts where one bar is slightly longer than another.
I’m not saying those differences aren’t real.
I just don’t feel most of them in my workflow.
Maybe one model gets me 72% of the way there and another gets me to 76%. Fine. Neither result is leaving my desk.
I’m going to read it, question it and check the sources. I’ll add things the model doesn’t know, delete things I don’t like, challenge assumptions, rewrite parts myself and probably send it through another iteration.
I’m not expecting the model to reach 100%. That’s my job.
So the benchmark question I care about isn’t simply whether one quant or model scored a point higher.
It’s whether that difference survives the workflow.
If one model saves me significant time, I care. If it consistently catches something the other misses, I care. If it can solve a problem the smaller one simply can’t, obviously I care.
But if both routes lead me to essentially the same final result after I’ve done my part, that extra percentage point has very little value to me.
The model isn’t the thing I’m optimizing.
The human-AI collaboration is.
Intelligence has a bill
There’s another reason this matters, especially once we move from my little lab to companies deploying AI at scale.
Someone has to pay for all this intelligence.
If you’ve watched a reasoning model work through a difficult problem, you know how strange it can look. It tries something, changes direction, reconsiders the request, starts another approach, realizes something is missing, searches, comes back and tries again.
Sometimes it feels like watching a machine gun of trial and error.
It’s genuinely impressive.
It’s also computation.
At the scale of one conversation, who cares? At the scale of thousands of employees and millions of interactions, I think you start caring quite a lot.
Especially because not all of that reasoning is necessarily solving a difficult problem. Sometimes the model is trying to figure out something we forgot to tell it.
Maybe the prompt was vague. Maybe the objective wasn’t clear. Maybe it didn’t have the right document. Maybe it misunderstood what we wanted.
A sufficiently capable model can often work its way through that uncertainty.
But we’re paying for it to do so.
How much AI reasoning are we paying for because we didn’t provide enough human reasoning upfront?
That question gets more interesting to me every time models become better at thinking for longer.
Sometimes I am the missing token
My smaller models get lost too. Obviously.
The difference is that I usually don’t let them stay lost for very long.
I’m watching.
At 20 tokens per second, I can see the wrong turn happening. Sometimes I’ll read two sentences and already know what’s missing.
I stop it.
“No, that’s not what I meant.”
Or:
“You’re missing this piece.”
And off we go again.
I don’t need the model to spend thousands of tokens trying to infer something that is already sitting in my head. I can just tell it.
Sometimes I am the missing token.
This is something I think gets lost when we talk about humans in the loop. The human isn’t only there to review the final output and make sure the AI didn’t screw up.
The human can be useful during the thinking.
I can prevent wasted reasoning, add information exactly when it becomes relevant, or recognize a dead end before the model does.
That’s not necessarily overhead.
Sometimes that’s the most efficient part of the system.
The cheapest token is the one you never generate
This makes me wonder whether we’re looking at AI efficiency too narrowly.
We talk about tokens per second, memory bandwidth, quantization, parameter counts, context length and cost per million tokens. All useful things to measure.
But what about the amount of thinking we made the machine do unnecessarily?
If I can replace 5,000 tokens of exploration with one sentence of context, that’s an optimization too. The same is true if a curated knowledge base saves a search loop, or a clear instruction stops the model from exploring three approaches I didn’t want in the first place.
Maybe we need to think about cognitive efficiency alongside computational efficiency.
Not “How much can this model think?”
How much thinking did we actually need to get somewhere useful?
Once you look at it that way, human participation doesn’t automatically look like an obstacle to automation.
It can be part of the optimization.
My little curated worlds
This is probably why I enjoy building small knowledge environments around local models so much.
I can create a curated wiki around a subject, add primary sources, transcripts and previous research, remove things I don’t trust, and tell the model what assumptions it can make or what it should question.
I don’t need it to know everything.
I need it to know this world.
Once I build that environment, the model itself starts feeling less important.
Today I’m using Qwen. Tomorrow it’ll be something else.
The knowledge stays. The instructions stay. The workflow stays. My point of view stays.
The model is one component.
A very important one, obviously.
But still one component.
Size and autonomy are different problems
There’s one distinction I need to make here, because otherwise this starts sounding like an argument against large models.
It isn’t.
Model capability and model autonomy are different things.
I could take a frontier model and use it exactly the way I use my local models: one exchange at a time, watching what it does, interrupting it, adding context and steering the work myself.
Or I could take a much smaller model, give it tools and memory, drop it into an agent loop and let it spend an hour trying to figure out a vague objective on its own.
The second system has the smaller model.
It’s also the one I’ve delegated more to.
That’s the part I care about.
Size tells me something about what a model can do. Autonomy is about how much of the process I choose to hand over to it.
I don’t have a philosophical problem with large models. I’d happily collaborate with the smartest model available.
What I’m less interested in is using intelligence as a reason for me to disappear from the process.
Once you separate capability from autonomy, the problem shifts. It stops being about what the model can do and starts being about how you design the space the model operates in.
I’m not just making an analogy here. This is instructional design.
Or, as I wrote recently: AI needs teachers too.
This IS instructional design
In learning and development, we don’t take everything humanity knows about a topic, dump it on someone and call that learning.
We decide what matters. We remove what doesn’t. We sequence things, provide examples when they’re useful, and build opportunities to try something, fail, adjust and try again.
We design the environment around the thinking we want to happen.
What does this model actually need to know? Which sources matter? What assumptions should it challenge? What am I bringing to the conversation? Where does the model’s responsibility stop and mine begin?
And underneath all of those questions:
Am I designing this interaction for collaboration, or for delegation?
Because those are different systems, and they need different things from AI.
I think I finally understand the 20 tokens per second thing
The industry seems to be moving toward AI that needs less and less from us.
Give the agent an objective. Come back later. Here’s your finished artifact.
I understand why people want that.
I’m just not sure I do.
I don’t want AI to become intelligent enough that I can stop thinking. I want AI that helps me think better.
I want to question it and have it question me. I want to bring another source halfway through because something it said reminded me of it. I want to catch a bad assumption before it becomes the foundation for the next 10,000 tokens.
I want to change my mind.
And I want to remain responsible for the final result.
Which brings me back to that little local model slowly printing words onto my screen.
Maybe the reason I like 20 tokens per second has very little to do with speed. It gives me enough time to stay in the conversation, notice when something is going wrong, and add myself back into the process.
So perhaps the question isn’t whether a 9B model can beat a frontier model.
It can’t.
That’s not the competition I’m interested in.
The question is how much model I need for the kind of relationship I want to have with AI.
If I want to delegate my thinking, then yes, give me the biggest brain money can buy.
If I want to collaborate, the equation changes.
I don’t need the model to do all the thinking. I need it to make my thinking better. And for that, 20 tokens per second is still plenty.





