Automate Your Agency

AI Context Windows 101

Alane Boyd & Micah Johnson Season 2 Episode 109

Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.

0:00 | 16:40

Send us Fan Mail

Free Resources & Next Steps:

About the Episode:

AI context windows just hit two million tokens, and if you're stuffing yours full of information and hoping for the best, you're doing it wrong. Alane Boyd and Micah Johnson break down what a context window actually is, why bigger doesn't always mean better, and what it means for how you use Claude, ChatGPT, and Gemini.

If you've ever gotten frustrated that AI "lost the plot" mid-conversation, or started giving you answers that made no sense, the context window is almost certainly the culprit. Most business owners don't realize that when AI's working memory runs out, it starts guessing, and it doesn't even know it's happening.

In this episode, you'll learn:

  • What a context window actually is: AI's short-term memory, and why it forgets without warning
  • The desk analogy that finally makes context windows click, and why leaving space matters more than filling it
  • Why hallucinations happen when the context window overflows, and how to prevent them
  • How Claude Cowork's subagents spin up their own context windows to handle complex tasks
  • Why your CLAUDE.md file and folder system are the most underrated tools for context efficiency
  • Where context windows are headed for Claude, ChatGPT, and Gemini, and the mobile phone plan evolution that tells the whole story

If you're ready to stop blaming AI for bad outputs and start understanding how it actually works, press play now.

Learn More about our AI Leadership Workshop

Winning businesses aren’t just working harder, they understand how to use AI strategically. Sign up for the AI Leadership Workshop, a 3-hour expert-led mastermind designed to give leadership teams clarity, alignment, and a practical path to stay competitive.

Book a call to customize your workshop today!

This episode is brought to you by Biggest Goal.

Every quarter your team spends evaluating AI is a quarter your competitors spend shipping. Most leaders feel the pressure but get stuck between ignoring AI and getting it wrong. More tools and more demos won't fix it. What actually works is hands-on training for the people doing the work.

For more information about our AI education programs, visit biggestgoal.ai.

Free Resources & Upcoming Sessions:

Connect with Us:

Disclosure: Some links are affiliate links, meaning we may earn a commission at no extra cost to you. Thanks for supporting the podcast!

Alane Boyd (00:00)
Two million token context window. But what the heck is a context window and why is it so important when using AI?

Micah (00:11)
So Alane, here's how I know I'm a complete nerd. I I look at yeah, no, I asked you and you said yes. No, I look at headlines like that and I actually get excited, like, two million token context window now. All right, that's cool. And I can quickly do the math because it's just multiplication. That is double the one million token context window that we've been working in.

Alane Boyd (00:17)
You asked me.

Micah (00:38)
With Opus 4.8 and the new Fable has a million token context window. Some of the GPT models has a million token context window. So we look at two million and we go, wow, that's amazing. But yeah, this whole episode is not just how I'm a nerd. Everybody that listens to this probably already knew that.

Alane Boyd (00:48)
It's just complication.

Probably so. That really makes me laugh. Okay, so I remember years ago now, it's been years, when you even talked about context windows, and I'm like, nobody cares, Micah. Like this is not important. We just need to use AI. And you're like, Alane, it is very important when you are using AI because it the context window is what's giving you

the output that you're looking for based on what you're feeding it and how much room it has to give you a good output. And I was like, Well, that's interesting.

Micah (01:29)
Yeah, it's a little bit like hard drive space. So, you know, as we were all younger and kids, hard drive space was a lot smaller than it is now. And now we have, you know, we can put a terabyte in our pocket and it's no big deal. But that would have taken up a whole room when, we had original computers filling rooms and, all of these.

Alane Boyd (01:32)
Mm-hmm.

Micah (01:53)
Just drastically different technology than what we have now. And the context window, we're kind of watching this play out in the same way. It started really small. So, you know, some of the early versions of GPT had like 32,000 token context window. And that felt big. 32,000 tokens. Holy moly, look at all this stuff we can shove into AI.

Alane Boyd (02:15)
Yeah, I remember one of our decks had a hundred thousand and two hundred and fifty thousand, and we were updating them as they came up with the newest release.

Micah (02:24)
Yeah, so it's constantly growing. But what it is, is it's essentially short-term memory for AI. That's AI's short-term memory. And I'm sure you remember this a lot too. It's faded now with clients, but a lot of people used to come in and go, Well, why can't I just feed everything in my entire business to AI and have it run my business for me? We would get questions like that all the time. And

Alane Boyd (02:48)
All the time, yeah. It it has

improved, I will say that.

Micah (02:52)
Yes, people are figuring it out. And the reason that you can't do that is because of the context window. The context window is too small to fit all that information. And when that context window is full, it starts hallucinating, which is just basically AI guessing. And so, it is like a person with memory problems. And they don't even know that they're forgetting. AI doesn't know.

Alane Boyd (03:10)
Mm-hmm.

Micah (03:18)
That it's forgot the original stuff that you gave it, or some of the stuff in the middle. It goes out of that context window and it is gone.

Alane Boyd (03:22)
Mm-hmm.

Right. And there's no hope. And it's why people think that can't trust AI or AI lost the plot and they can't get it back to answer really what they want it to. And it really all revolves around this context window and that it does not have a true memory.

Micah (03:43)
Yeah. And some of it was just user error, right? Like if you overload a person with too much information and then ask them to recall a hundred percent, that is not going to happen for most people. But we don't give that same leniency to this technology. We say, Well, we gave you all of this. Why can't you remember all of it? AI, you jerk.

Alane Boyd (03:53)
Mm-hmm.

We have a double standard for sure. So, Micah, I love the analogy that you came up with recently with the desk.

Micah (04:08)
Yes.

Yes. Okay. So the way that I think paints a really good picture mentally is if you imagine a desk and the size of the desk is the context window. I even look at the context window and I gotta stop myself from doing this thinking a million token context window or a two million token context window means I can load a million

Tokens worth of information into AI and then have it work seamlessly and flawlessly and perfectly. But it's not the million tokens is your entire desk space. And if you cover your entire desk with papers and folders and equipment and computers and monitors, you have no more room to actually do the work. And the context window is exactly the same concept. If you fill it full, there's no room.

Alane Boyd (04:57)
Hmm.

Micah (05:05)
To do the work. The AI doesn't have any room to process anything.

Alane Boyd (05:08)
I loved that analogy too, because it really helps you understand that there's the input and the output that you're working with with the token window. And a lot of times it's not a one and done thing. You don't just have AI do something and then you're good to go. A lot of times it's a conversation that you might be having with it, fine-tune things, tweak things. Well, that's all part of it.

Micah (05:28)
Yeah, that all has to fit in the context window and the stuff that happens behind the scenes. So when you're using tools like Cowork and it goes thinking, still thinking, almost done thinking, right? All of that as it's processing, guess what? That's using the context window. And if there's no room in the context window, there's no room for it to think. And while it's trying to process, it's actually kicking old stuff out of the context window. It doesn't even know it's happening.

Alane Boyd (05:56)
Yeah. And different models that you might use, like within Claude or within ChatGPT also have different context windows or what they're able to do within those context windows.

Micah (06:09)
Yes, absolutely. I mean, Sonnet 5 just came out for Claude and it introduced a million token context window for Sonnet, the the mid-tier of the model. but before that, Sonnet was 200,000 tokens and it capped at that. And so if you think about all the stuff that you're using AI for throughout the day, and especially with Cowork.

And definitely with code as well. But if we think of Cowork, right, you're loading a CLAUDE.md file if you have your file system set up correctly. You're loading additional context in as you're working. It might be doing web search for you. It might pull all this stuff in. And a 200,000 token context window all of a sudden feels really, really small when you've got to keep about half of that open for workspace.

Alane Boyd (06:55)
Uh-huh.

And I mean it goes back, we talk about this on so many episodes, but having your folder system set up with a CLAUDE.md file, that will help so much with the context window and the usage that you're doing because Claude, if you're using Claude, knows exactly where to go to find that information. And so you have some more wiggle room in there for your output.

Micah (07:21)
Yeah. And I I would say, you know, we talk about Claude a lot because it is doing spectacular things. It is extraordinary. OpenAI's Codex is similar to the Claude desktop. I wouldn't say it's one to one yet. but it also has the ability to access your file system. So it's getting closer. Google with the announcement of Gemini and the two million token context window, you're right. That is a hard thing to say. That is a tongue twister.

Alane Boyd (07:27)
Mm-hmm.

Mm-hmm.

Micah (07:51)
That they're also introducing a desktop app. So the other players in this space and the other providers, they are all going the Cowork and the Claude desktop type direction. So if you're not using your file system now, I'm kind of echoing what you're saying, Alane, but bringing it to the whole AI world, eventually you'll have to, because that is where all this is going. And it's a huge boon to the context window approach, like you're saying. If you just use chat, you have to load.

Alane Boyd (08:09)
Mm-hmm.

Micah (08:19)
Everything in at once. That eats up the context window. Your chat eats up the context window. It's thinking eats up the context window. And pretty soon you're not getting the good results that you need to get out of this. but when you have the file system access, it can load a little piece and load another file. It doesn't have to load the entire file system in just to understand what you're trying to do.

Alane Boyd (08:21)
Mm-hmm.

Mm-hmm. And then, you know, what we're really doing a lot with the Skills and Cowork and using Cowork generally is they've made it easy for us to use agents essentially. We're just using very easy way of using agents. I don't know a good way to say that, but that's what's happening behind the scenes when you're using Cowork, or these agents. And those agents, go ahead, Micah.

Micah (09:07)
I was just gonna say Cowork behind the scenes is running agents and sometimes agents that are running other agents.

Alane Boyd (09:14)
Yeah, think about what we used to have to build as an agent now can be done within Cowork, because it is an agent working behind the scene. They just made it wrapped up really nicely for everyday people. But those agents, when you have them working, they need room to work. They're gonna be using that context window.

Micah (09:32)
Yeah. Let me throw something out there though, because this is where it gets really intriguing on the context window. If you're using a tool like Claude Cowork, and let's say you're running Sonnet 5 as your model, you have a million token context window. Your main conversation is in that million token context window, but Claude Cowork is strong enough to say things like, and when you work with it, you'll see it articulate this and essentially say,

Let me start up some subagents so they have their own context windows and it's not, you know, filling up my main context window. And so Claude even has the strategy of saying, hey, if I start up three other agents to help me do this, then they each have a million token context window for each of them. They can process, pull in information, summarize it, and do an output that goes

to the main context window. And now the main context window only needs a little bit more taken up because it's all happening in those subagents.

Alane Boyd (10:37)
That is pretty cool, Micah. I have not paid that close of attention when it's giving me the feed of what it's doing. So I appreciate that.

Micah (10:45)
Yeah, you'll see it it has a new interface in Cowork that'll come up and it'll actually show you like miniature boxes that represent they don't look it's just boxes, right? And it'll have a pulsing circle and like that's the agents working behind the scenes. But you'll see a string of three or six subagents come up and you'll actually be able to see them work and then some of them will start getting check marks, and that means their job is done. Yeah.

Alane Boyd (10:49)
Hmm.

Mm.

Hmm.

Hello.

that's really cool.

that gives it that more of human feel too. it's cool that you can see what's happening, but also it's like these little robot people are doing things and they're getting things done.

Micah (11:23)
Yeah, and all of this, these are all strategies that are trying to combat the limitations of context windows. So even though we now have two million context windows in Google models and the latest version with ChatGPT 5.6 has limited availability, but it's a 2 million token context window. We just have to remember that

Eventually we won't worry about context windows. Like we don't worry about hard drive space too much, but I think it's gonna follow the same kind of track. Even though we have more hard drive space than ever, we also introduced higher video rate formats that are giant like all the stuff that we do takes up more space. So we need more hard drive space. And we see the same thing with AI and the processing. The more that we do, well, it needs more context.

Alane Boyd (12:11)
No.

Micah (12:16)
So that works hand in hand. The context windows are getting bigger, but then we're using Opus and Fable, which uses five times the amount of context window to do the same task. So it's not like the context windows are getting bigger and we just have massively more amount of space. following that same trajectory. We still have to be mindful, at least today, of

context windows and the space that we have available in the context windows, but at least we have these other strategies like the file system and subagents that can be created to handle all of that.

Alane Boyd (12:50)
So I have an analogy for you. Well, I I don't know if it's actually an analogy, but you know what this exactly reminds me of? And we keep hearing that AI is going to be a utility cost, just like we pay for electricity and water and stuff like that. Well, this reminds me exactly of when mobile phones came out. At first, we paid for minutes, then we paid per text. Remember, it was like,

Micah (12:53)
Okay.

Alane Boyd (13:19)
20 cents per text message or something like that. And then you got so much network to internet time. And then you'd have to pay for overages. And then ATT, I think, came out with unlimited. And then they backtracked and they went back to so much. And now it's now you get unlimited minutes, unlimited texting, and unlimited internet or whatever you want to call it, data. And so I see that kind of happening here is where we're gonna see this.

Micah (13:20)
yeah.

Alane Boyd (13:48)
with context windows, what we have accessible, where it's kind of a battle between the AI companies on what they're gonna offer with the context windows, what the pricing's gonna be, what usage is like, and eventually it'll just be a standard price.

Micah (14:03)
Yeah, I I think you're really accurate on that. I can't wait to see like Verizon style and ATT style commercials, but for AI. That's the mint mobile ones where they're funny.

Alane Boyd (14:13)
I'd like to see mint mobile one types of commercials.

Micah (14:18)
Yeah. So to kind of wrap up this whole thing about context, the the big takeaway here is even though the context windows are getting bigger, you still have to be mindful about what you're putting into them. If you're using tools like Cowork, you have the additional strategies that Cowork can already benefit with the context windows. But the main takeaway is just don't read million token context window, two million token context window and go, great, now I can completely fill it with

All of this information and still think that AI is going to solve all these problems for you. It's not, you still have to make judgment calls. You're still the human. You're still in control. You're still guiding it. But to get the best results, don't make the user error mistake by putting too much information in there and then getting mad at AI.

Alane Boyd (15:06)
Good one. And if you haven't yet, get your folder system set up, get your CLAUDE.md file set up, go to your.biggestgoal.ai. We have a free Cowork Masterclass where we walk you through exactly what to do. So there's zero excuses.