I Built A "Cognitive Harness" That Upgrades Your AI And Makes You Smarter
Your Brain Can Run 3 Mental Models at Once. Your AI Can Run 300.
Editorial Note
I’ve spent 10 years creating the most comprehensive encyclopedia of mental models ever. It includes 2,000+ mental models drawn from every major discipline and industry relevant to career success, business success, investing, personal growth, parenting, health, and relationships.
That wasn’t the real challenge though.
No matter how many mental models I gathered, applying the right combination of them in the moment remained elusive. The brain can’t recall 2,000 mental models, pick the 5-10 most relevant ones, and then fire them off in sequence. All in just a few seconds. That would take hours, which makes the entire enterprise impractical.
Thus, most people, even those who’ve studied mental models, still end up in a cognitive rut, using the same handful of models over and over.
For all of human history, we’ve had to settle for two modes of thinking:
System I (fast, automatic)
System II (slow, deliberate)
What’s different about this moment is that AI can fill in the gaps where the human brain is limited. In so doing, it can magnify our cognitive abilities, effectively increasing our IQ. This human + AI thinking system marks the dawn of System III thinking, which is both fast and deliberate.
I’ve spent the last year exploring how to deploy AI to this end. The end result was an AI skill that fires every time you use AI. This skill makes your AI smarter while teaching you more mental models.
I shared version 1 of the skill in February. It got more positive feedback than any prompt or skill I’ve ever shared before. So, I decided to create a version 2 of it. That’s what this monster article is about.
What You Get Today
Free Subscribers
The top lessons I’ve learned on my 10-year mental model journey, compressed into 10,000 words.
It includes:
How version 2 of the AI skill works
The key milestones along my personal journey
A survey of the most important academic studies on mental models ever done
A survey of some of the most important cognitive harness case studies
A survey of how AI pioneers are using mental models to make smarter decisions
How Human+AI thinking systems completely change the potential of cognition
Paid Subscribers
Version 2 of the mental model skill, which includes:
A version for AI chat, and a version for agentic AI
The full 2,000+ mental model encyclopedia
An encyclopedia of 400 paradigms
Step-by-step install instructions, with screenshots
A six-step router: stuckness diagnosis → worldview identification → 338 rival worldview pairings → question de-biasing → 61 thinking recipes → self-verification
In addition to this skill, you get $2,000+ in other paid perks (books, skills, classes, mental model manuals).
FULL ARTICLE
The first time I was told I wasn’t smart, I was eight years old.
I had missed the cutoff for my school’s gifted and talented program by a single IQ point.
At the time, all I knew was this:
IQ is a fixed quality that you can’t change.
Your IQ predicts much of your career performance.
My two best friends got into the program, and I didn’t.
So there I was faced with a horrible feeling of inadequacy along with two truths that shouldn’t be able to co-exist:
The belief that I could do anything (given to me by an immigrant mom).
The message from the outside world that I would always be limited.
That single number going into third grade felt like a life sentence.
Psychologists have a name for what a near-miss like this does to the human psyche. Research finds that falling just short of an important threshold stings sharper, and lingers longer, than missing by a wide margin. For me, that sting didn’t fade for decades. The “almost” turned into fuel: if I couldn’t be born smart, I would outwork everyone who was.
When I Was 16, I Found A Shortcut To Effectively Increase My IQ
The shortcut was reading.
At the time, my business partner, Cal Newport, and I had co-founded a web development company despite having no business experience and little web development experience.
To compensate for this, we took the first $1,000 of our earnings and bought tons of books.
It worked amazingly well!
For just $15, we could get somebody’s life lessons distilled in 200 pages. We started with Adobe Photoshop, HTML, business, typography, and color theory. And, the client ended up being very happy with what we created.
This learning shortcut turned into a multi-decade habit of ravenous reading. I still think that the information in books has the highest concentration of wisdom with the least distraction.
The next shortcut came in 2015 when…
I Found Out How To Exponentially Increase My Rate Of Learning
At the time, I had a thought that changed everything:
“If I’m going to read thousands of books in my life, why don’t I just deliberately learn how to read and learn faster and better?”
So, I took a step back and thought more deeply and systematically about intelligence amplification:
I spent 100+ hours reviewing the most rigorous academic research on IQ intervention.
I collected all the learning models I could find across different fields.
I found research showing that diet and exercise interventions can have a profound impact on intelligence (see the appendix for the most fascinating research).
I studied the learning habits of top entrepreneurs, investors, and executives.
As I was doing all of this, one thought was guiding me:
“Machines like the horse-drawn plough reduced the importance of physical strength in the world. What is the plough that augments the mind and reduces the relative importance of IQ?”
As an adult, this question felt significant to me. My logic was this:
Intelligence is important across humans, animals, and systems.
If intelligence were malleable, it would have profound implications not only for my own life but also for democratizing access to its benefits in society.
This period led me to several ways to learn faster. Below are a few of the simplest and most profound takeaways packaged into articles I wrote:
Modern Polymath: People Who Have “Too Many Interests” Are More Likely To Be Successful According To Research (92,000 likes)
The 5-Hour Rule: If you’re not spending 5 hours per week learning, you’re being irresponsible (68,000 likes)
The Learning Loop: Memory & Learning Breakthrough: It Turns Out That The Ancients Were Right (2,600 likes)
Fractal Reading: Augmented Reading: Learn 10x Faster And Better With AI
The Explanation Effect: Memory & Learning Breakthrough—It Turns Out That The Ancients Were Right (10,000 likes)
How Elon Musk Learns Faster And Better Than Everyone Else (21,000 likes)
Over time, my opinion on IQ as a fixed, all-important trait evolved, and I concluded...
Forget About High IQ. What Really Matters Is Augmented Intelligence.
My experiences and research have convinced me to adopt a growth mindset about intelligence. Many proven interventions increase or augment IQ. Not only that, the IQ of people in developed countries has increased dramatically in the 20th century, showing that it’s not a fixed variable...
At the same time, I learned that the line doesn’t just go one way. Over the last two decades, the trend has reversed itself:
Then, I came across a fascinating pattern...
The World’s Smartest Entrepreneurs And Investors Use Mental Models To Augment Their Intelligence
I learned about the power of mental models when I began studying the learning habits of many of the world’s most successful entrepreneurs, investors, and innovators, including Jeff Bezos, Elon Musk, Warren Buffett, Peter Thiel, Howard Marks, Ray Dalio, Drew Houston, and many more. Each had the peculiar habit of collecting and coining mental models.
Some went even further. For example, Charlie Munger, Warren Buffett’s lifelong business partner, built an entire operating system for making smart decisions with a “latticework of mental models.”
Here’s what he did:
Identified that mental models are the most powerful unit of learning. These models are compact, reusable ideas that explain how important parts of the world work.
Collected the big models from the big fields. Each field has unique insights. And not all models are created equal. Some are more important than others.
Learned the most important 100 mental models. Enough to cover the most important models. He chose not to spend forever learning thousands of models.
Applied each of the models religiously. Every time he made an investment decision, he used two checklists.
At a more profound level, Munger believed reality was one connected system. That while academia splits reality into separate disciplines, reality itself does not respect the boundaries between economics, biology, and psychology. Said differently, there are deep, fundamental principles that apply across fields that most people overlook, because those principles are not taught in any one specialty.
Elon Musk echoed the same sentiment:
“It is important to view knowledge as sort of a semantic tree — make sure you understand the fundamental principles (Musk calls these ‘first principles’), i.e. the trunk and big branches, before you get into the leaves/details or there is nothing for them to hang onto.”
—Elon Musk
I created the following visual to make sense of the quote:
Because knowledge is interconnected in a tree-like structure, mental models can often be used across many fields to generate big, unique insights and make smarter decisions.
Munger returned to the power of this idea at the 2007 USC Law commencement:
“There are all these other things that you should know in addition to history. And those other things are the big ideas in all the other disciplines. And it doesn’t help you just to know them enough so you can prattle them back on an exam and get an A. You have to learn these things in such a way that they’re in a mental latticework in your head and you automatically use them for the rest of your life.
If you do that, I solemnly promise you that one day you’ll be walking down the street and you’ll look to your right and left and you’ll think, ‘My heavenly days, I’m now one of the few most competent people of my whole age cohort.’ If you don’t do it, many of the brightest of you will live in the middle ranks or in the shallows.”
—Charlie Munger
Sit with the implications of what Munger is claiming, because they operate on three levels at once:
The personal claim. The ceiling on your decisions is not your intelligence or your effort. It is the number of models you can actually reach for in the moment. Almost everyone reaches for the same three or four.
The professional claim. Across many fields, the people who carry the best models from many disciplines beat the specialists who run everything through one. Research backs this claim. Consider the 20 most significant scientists in history, as ranked in Human Accomplishment, which scored thinkers by how much space reference works devote to them. 15 of them were polymaths. Newton. Galileo. Aristotle. Kepler. Descartes. Huygens. Laplace. Faraday. Pasteur. Ptolemy. Hooke. Leibniz. Euler. Darwin. Maxwell. All polymaths. Not only that, the founders of the largest companies in the world are nearly universally polymaths. Bill Gates, Steve Jobs, Warren Buffett, Larry Page, Elon Musk, Mark Zuckerberg, Jeff Bezos.
The civilizational claim. The real limit on our breakthroughs has rarely been how much we knew. It’s been how few minds could combine what they knew across the walls between fields. That kind of mind, holding many disciplines in view at once, has been one of the rarest advantages in history, and building such a mind costs a lifetime of work.
Munger isn’t alone.
Many entrepreneurs have their own unique approach to collecting, creating, and chaining mental models...
Everyone Who Thinks For A Living Collects And Coins Mental Models
Elon Musk is famous for reasoning from first principles instead of by analogy, a concept that is common in physics:
Source: Kevin Rose Podcast
He’s also famous for using other models, such as thinking in the limit, probabilistic thinking, attacking the bottleneck, and his own 5-step iterative product development algorithm.
Ray Dalio, who built the largest hedge fund in the world, wrote 600 pages of Principles based on the same instinct:
“Principles are concepts that can be applied over and over again in similar circumstances, as distinct from narrow answers to specific questions. Those who understand more of them, and understand them well, know how to interact with the world more effectively than those who know fewer of them or know them less well.”
Ray Dalio, Principles
Dalio uses the word “principles” where Munger says “models,” but he is pointing at the same thing.
Brian Chesky, the co-founder and CEO of Airbnb, coined his own mental model, the 10-star experience, and built his company around it:
John and Patrick Collison, the founders of Stripe (≈$150B valuation), were so taken with Munger’s latticework that they republished his book, Poor Charlie’s Almanack, under their own publishing press.
The Nobel laureate Herbert Simon, one of the pioneers of artificial intelligence, put it plainly decades ago:
“The quality of our mental models determines how well we function in the natural world.”
The list of unique models people regularly use is surprisingly varied:
…and the list goes on.
Mental models were just as impactful in my own life…
How Mental Models Changed My Life
Explaining the massive shift that mental models made in my life, I wrote the following about my firsthand experience in Most People Think This Is A Smart Habit, But It’s Actually Brain-Damaging (2020):
I saw reality on a much deeper level — and on a fundamentally different level. I looked back on many of my old mistakes, and thought to myself, “Oh my God! If I had only known this or that mental model…” I wasn’t just learning new strategies or hacks.
On some level, I could relate to some of my favorite movie characters just after their intelligence had exploded…
First, I saved over $100,000 dollars…
For example, 12 years ago, I borrowed $100,000 from friends, vendors, and banks (at high interest rates) to keep a struggling website I created alive. Rather than facing obvious indicators that the idea wasn’t working, I kept on doubling down. I was in love with my idea, and I didn’t want to admit defeat. The company died a slow, painful death.
Now that I understand mental models, I see how one cognitive fallacy — “sunk cost fallacy” — caused my poor decision-making. This mental model has helped me discontinue failing ideas much more quickly.
Today, when I consider new business ideas, instead of just imagining how great they’re going to be, I also envision what could go wrong — the inversion mental model — saving a lot of my time and money upfront. For example, a few months ago I had the idea to create a book summary of the month club where I would write a weekly in-depth summary of a life-changing book. I got really excited about it and spent 10 hours thinking about how great it was going to be. A few days of planning later, I decided to take a step back and honestly assess the downsides. Almost immediately, I started seeing some glaring roadblocks, and the new shiny object was no longer as exciting. Soon after, I decided to just focus on our core businesses. I’m very happy that we did. Twelve years ago, I would have jumped in straight away.
Also, our article virality shot through the roof…
The success of our business is directly related to the number of views each article gets. So, being able to get hundreds of thousands of visitors per month without paying for ads is a big deal.
After learning about the 80/20 Rule (which is now one of my favorite mental models), I started asking the most successful article writers what the 20% activities were that give them 80% of the results. Almost all of them mentioned that titles were key. Previously, I viewed the titles as just an afterthought.
Because of this insight, my team and I restructured our entire article creation process:
Now, we create titles before we write articles.
Rather than spending 5 minutes on titles per article, we dedicate 5–10 hours per article.
For every article, we brainstorm over 30 titles rather than a handful.
We test the article titles before we publish them.
Our team has now spent over 1,000 hours studying the patterns of titles and testing nearly 5,000 titles. As a result, we have a fundamentally different and better understanding of what makes articles go viral.
This is one of the mental models I use for my articles to be viewed tens of millions of times…with the average article now being viewed 150,000+ times.
Finally, I started making a lot more money.
I have had people hire me for six-figure consulting contracts.
I began charging $500 per hour as a coach and consultant.
I even got a five-figure speaking gig from an article.
When I launched a new company, we were able to create a business strategy based on “breakthrough knowledge” that is hard to replicate. This one strategy alone made us over 7-figures in revenue.
I once heard a coach talk about changing a client’s way of seeing the world in a way that would blow their mind. When he looked into his client’s eyes and could see him or her really getting it, he’d say, “Now, you’re in my reality!”
That’s how I felt.
Reality somehow feels different on an aesthetic level — as if I’m cutting through the levels of illusion and noise we normally see and getting a more direct view.
The best way I can describe this is that it’s like wearing augmented reality glasses that constantly feed you relevant wisdom about the situation you’re in.
Then, around 2021, I started to feel limited exploring mental models just through the lens of career success.
Mental Models Are More Than About Career Success. They Are About Life Success And Happiness.
Around 2020 and beyond, two things changed:
Society changed. I realized that entrepreneurs that I had long-respected weren’t the best role models outside of business. I won’t name names, but you know who I’m talking about.
I changed. As a parent of two children at the time and in a 20+ year marriage, although I loved business, it was no longer the center of my life like it was in my 20s. My family, quality of life, health, impact, and personal growth had taken center stage.
These two changes led me to explore mental models for:
Relationships
Health
Emotional Regulation
Developmental Growth
Somatic Awareness
Etc.
This had a profound impact on me. Here’s an example of one shift.
When I first became a parent at 26, I was flying blind. None of my friends had children yet, and the books about parenting for dads were woefully inadequate.
Then, last November, 17 years later, my wife and I had our third child—Naya:
That’s when I really saw the shift in myself and the power of mental models.
Here’s a concrete, recent example. With my first two children:
Awareness. I was only focusing on major milestones (eg, walking, crawling, talking), and I ignored smaller, real-time milestones. This happened because my model of development was low resolution. This made it hard for me to see how much my kids were evolving every single day.
Knowledge Transfer. Parenthood and business felt like completely different domains. It was hard to transfer commonalities between them. Therefore, it was hard for me to apply parenting lessons to life/business and vice versa.
Parenting Mental Models. I wasn’t aware of mental models for emotional regulation techniques or somatic awareness. This made it hard for me to constructively process the most difficult parts of parenthood. Which led to overwhelm.
Now things are different on multiple levels:
Awareness. Rather than focusing on big milestones, I’m able to notice the smallest ones. This makes the experience feel more present and alive.
Knowledge Transfer. Now that I’m able to abstract experiences into deeper models, it’s easier for me to learn something in one area of my life and generalize it across other areas. Thus, each area of my life feels more integrated. Parenting feels related to every other area of my life, which makes it more engaging.
Parenting Mental Models. I feel well-resourced as a parent because I have a whole toolkit of mental models that help me. Some come from the field of parenting. Others are transferred from other fields.
And now that AI exists, I’m able to take things to another level.
Given how deeply I’ve studied learning, growth, and developmental psychology, I have many mental models in that area.
Earlier in my daughter’s development, I noticed that she learned to intentionally let go of things she was holding on to. Normally, I would think nothing of it. But, I now see letting go as an important developmental milestone. And I’m able to relate to it by thinking about how it applies in other domains (eg, emotionally letting go) in my life. Seeing her learn to let go inspired me to do the same better in other areas of my life.
Seeing this pattern, I was then able to ask AI the following question:
“Many people notice milestones like standing, but they miss milestones like falling. Or they notice grabbing, but not letting what go. What are other milestones of development that often go overlooked? And how do those milestones mature as humans evolve, and how are they relevant to me now?”
This yielded a table with 68 developmental pairs. You can see the full list in the appendix, but here’s a sampling:
This table has shifted how I think about development on a deeper level and led to deeper conversations with AI.
My point is that when I was younger, I would’ve completely missed noticing almost all of my kid’s milestones. Now, I’m not only able to experience them in real-time with her, but I’m also able to participate in and learn from them, because I have a deep understanding of mental models and AI.
These shifts transform parenthood in the best ways. They deepen my relationship with her. They help me grow, and they help me help her grow. And they make parenthood more fun.
From Anecdotes to Evidence:
At this point, you could object. You could say that this list and my experience are survivorship bias:
That successful people are just a small sample.
That these models are just the story they tell about their luck.
Fair.
So set the anecdotes aside and look at what happens under controlled conditions...
Seven Studies On Why Range Beats Brilliance
Range is carrying many models from many disciplines, and reaching for the right one when it counts. Across seven studies and very different methods, the same finding recurs: expertise is built from models, and the breadth to deploy the right one beats raw brilliance:
Experts think about what mental models are relevant to a problem before trying to solve it. Amateurs don’t.
Experts with more, diverse mental models outperform experts with few models from one field.
The most creative experts have the widest range of interests.
Your range of outside interests predicts success better than your IQ.
Storage of mental models is not enough. Retrieval is key.
The basis of expertise is the number of stored mental models.
The pattern applies to teams as well.
These aren't anecdotal claims. Each one comes from research spanning four decades: a 20-year forecasting tournament tracking 28,000 predictions, creativity studies following scientists over 25 years, controlled experiments in analogical reasoning, and the original chess expertise research that helped launch cognitive science.
Together, they point to the same conclusion: the ceiling on your thinking is not just your raw intelligence. It’s the number of mental models you can store, retrieve, and fire at the right moment. If you want the full breakdown of each study, including the methods and specific results, I've included it in the appendix at the end of this article.
Bottom line
For decades, collecting mental models was the best advice anyone has given on how to think.
But, Charlie Munger is the one who turned it into an operating system.
To understand why what Munger did was so rare, you have to first look at the more ordinary things people do with models…
There Are 5 Levels Of Mental Model Mastery. Munger Was One Of A Few People To Reach Level 5.
There are five levels of engagement with mental models, each rarer than the last. Munger reached the fifth, and until very recently, very few people could follow him there.
Level #1 - Using Mental Models Without Knowing It: Close your eyes right now. If I asked you to walk outside of the building you’re in right now, you could probably do this. The reason you can is that your mind unconsciously built a model of the building in your head that you can use to navigate physical space without sight. Your mind automatically builds all sorts of other models. Models of individual people, models of products, models of businesses, models of society.
Level #2 - Using Models On Purpose: Not all models are helpful. Some are so outdated or overused that they do more harm than good. This is why using models consciously is so important. For example, you use one consciously the moment you catch yourself thinking ‘that’s a sunk cost’ before finishing a book you stopped enjoying 100 pages ago. You know the model by name, and you reach for it on purpose.
Level #3 - Collecting the Best Models: The most successful thinkers and experts don’t just use models based on their life experience. They consciously collect and use the best ones of all-time. For example, the large list of successful model collectors I shared at the beginning of the article.
Level #4 - Creating Your Own Models: Jeff Bezos didn’t borrow regret minimization, he coined it. Nassim Taleb gave us antifragility. Ray Dalio built believability-weighted decision-making. Annie Duke named resulting.
Level #5 - Writing the Rules For Running Models: Munger did something no one had ever done before. He wrote down the rules for operating mental models.
Charlie Munger’s 5-Step Operating System For Thinking
A mental model is a compressed piece of how the world works, borrowed from a field that spent centuries figuring it out.
Supply and demand from economics
Compounding from mathematics
Natural selection from biology
Each one is a small machine for prediction. You feed a situation in one end and a better guess comes out the other.
No single model is enough, because every field only sees its own slice of reality. Stack enough of them across enough disciplines and you stop seeing a narrow problem and start seeing the whole one.
Here is his operating system, in his own words:
Rule #1: Learn Multiple Models
The first rule is that you’ve got to have multiple models—because if you just have one or two that you’re using, the nature of human psychology is such that you’ll torture reality so that it fits your models.
It’s like the old saying, ‘To the man with only a hammer, every problem looks like a nail.’ But that’s a perfectly disastrous way to think and a perfectly disastrous way to operate in the world.
Rule #2: Learn Multiple Models From Multiple Disciplines
And the models have to come from multiple disciplines — because all the wisdom of the world is not to be found in one little academic department.
Rule #3: Focus On Big Ideas From The Big Disciplines (20% Of Models Create 80% Of The Results)
You may say, ‘My God, this is already getting way too tough.’ But, fortunately, it isn’t that tough — because 80 or 90 important models will carry about 90% of the freight in making you a worldly-wise person. And, of those, only a mere handful really carry very heavy freight.
Rule #4: Use A Checklist To Ensure You’re Factoring in the Right Models
Use a checklist to be sure you get all of the main models.
How can smart people be wrong? Well, the answer is that they don’t…take all the main models from psychology and use them as a checklist in reviewing outcomes in complex systems.
Rule #5: Create Multiple Checklists And Use The Right One For The Situation
You need a different checklist and different mental models for different companies. I can never make it easy by saying, ‘Here are three things.’ You have to drive yourself to ingrain it in your head for the rest of your life.
The lattice Munger described, dozens of ideas firing together and checked against each other, is wonderful in theory.
But over time, as I tried to implement it, I noticed diminishing returns: each new model I learned added less value than the one before...
6 Reasons The Latticework Ultimately Lets Down Every Human Who Tries It
I have been studying mental models for years, and I have never met a single person who actually runs the latticework the way Munger described it. Not one.
Here are the six reasons why:
Invisibility of our thinking. You cannot see the unconscious models already running your decisions, so you cannot correct them.
Limited time. You cannot learn and operationalize all the good ones unless thinking is your full-time job.
Forgetting models. Even the models you have learned rarely surface at the moment they would help, which is the whole reason Munger reached for written checklists.
Need for immediate decisions. A checklist works when you have time to deliberate. It does nothing in the middle of an argument or when a fast decision is required.
Limited working memory. Your working memory holds only about three or four things at once, so you can never actually run many models in parallel.
Emotions distorting thinking. Knowing a model is not the same as applying it when you’re emotionally triggered, and the bigger the decision, the more feeling floods in. You can recite sunk cost and still refuse to walk away from what you poured years into, or know the priority cold and still not bring yourself to say no.
In practice, almost everyone collapses back to the same three or four models, over and over.
And they’re not even choosing those three or four because they fit the problem. They reach for them because they are familiar, comfortable, already wired to fire. Picture the person who runs everything through incentives, supply and demand, and first principles, while the biologist’s lens, the historian’s, and the negotiator’s sit unused on the shelf, even when one of them would have cracked the problem wide open.
Seeing a situation through a model the way those experts do is not a matter of reading a definition once. It takes dozens of hours to learn a single model well enough to fire it automatically wherever it applies. That is the price of a single model. And the number you need is far larger than anyone has admitted.
Munger thought 80 or 90 would cover it. The models genuinely worth having run into the thousands. Do the arithmetic: thousands of models at dozens of hours each is more time than any lifetime holds.
I should know.
I spent five years cataloging them, my encyclopedia is past 2,000, and I am nowhere near done. Everyone who chases the latticework, including the people who admire it most, learns a sliver, and truly operationalizes a sliver of that sliver.
Researcher on polymathy Dean Keith Simonton puts the challenge well:
“Someone can’t just say, “Well, as of today I’ll have extremely broad interests so that tomorrow I’ll be a creative genius.” Doesn’t work that way. The broad interests are involved in expanding a person’s knowledge base, and that takes considerable time.”
—Dean Keith Simonton
I had a head start most people never get. For four years, I published a 10,000-word manual on one mental model every month. I built an encyclopedia. And still, on a normal Tuesday, with a real decision in front of me and no time to deliberate, I reached for my same few favorites. I was using more models than almost anyone I knew, and I still felt like I was living at 1% of what the latticework promised.
When I talked to mental model students about this, the story got worse, not better. They saw the value immediately. Then they saw the work, the years of study, the thousands of models, the checklists you build and rebuild, and they quietly decided it is not for them. They were not wrong about the cost. For a human, the cost is the whole problem.
So the latticework was always two things at once:
It was the most powerful idea anyone has had about how to think.
And it was almost completely out of reach for the average person.
Years of work bought you a sliver, and that sliver could never run at full bandwidth. The dream of mental models was a mind deep enough to hold a hundred models, fast enough to fire the right one on cue, and wide enough to run hundreds at once. No human ever had it.
It’s as if mental models are this incredibly powerful intelligence enhancer, but the human brain simply isn’t designed to run them at their max. It’s like receiving a technology from the future, but not having the right power source to make it run.
Bottom line:
While mental models are extremely powerful, it takes thousands of hours to truly master them. And even then, it’s simply hard for the human brain to use multiple mental models at once in real-time in all of the relevant situations in our life where they could make a profound difference
Then I added one sentence to the end of a question I asked AI, and the limit I had treated as a law of nature turned out to be surpassable...
One Sentence Brought Every Limit Down
Here is the thing that took me embarrassingly long to see. I had spent five years treating the six-limit ceiling as fundamental. It is not. Rather, it is a fact about how human thinking works. Think with AI, and the ceiling moves.
The sentence I added was, more or less, this:
“Break this down to first principles, then find the incentive everyone is ignoring, then run it through the lens of a biologist, then steelman the opposite of whatever you just concluded, then tell me what a historian would notice.”
The first time I did it, the answer was not a little better. It was a different kind of answer. And, it came back in seconds.
This made me realize that I could almost treat mental models as a command language and use AI to process multiple mental models in sequence.
To put the power of having a command language in context, I share the following in Vibe Prompting Method:
“Every field has this hidden vocabulary—a minimal set of high-leverage terms that compress expertise into commands. Musicians have it. Coders have it. Designers, writers, analysts—all fields have it.”
And once you figure out that command language, you can do incredible things. For example, in this 2-minute video, a musician uses words I’ve never heard before to compose an amazing beat:
In another video, we see famous movie director Martin Scorsese immediately take a scene and use his expertise to command AI in a way that a novice never could.
What I now see is that mental models are the command language for thinking. Point models at AI, and you can make it smarter.
This experience begged a question:
What if AI could operate Munger’s system?
What I now believe is that the six limits I mentioned earlier are human limits. Not one of them belongs to AI:
AI thinking is visible. You cannot see the models running in your own subconscious. With AI, you name a model, and it applies exactly that one, on command, and you can read the text of exactly how the AI is applying it.
AI already knows the encyclopedia. You cannot learn and operationalize them all in a lifetime. AI already has. Every framework I spent years writing up is sitting inside it, fully operational, for the price of a monthly subscription.
AI doesn’t forget. The right model rarely surfaces when you need it. The machine forgets nothing. Its whole library can be loaded, so the model that fits is always there, no checklist required.
AI thinks in seconds. A checklist is useless in the heat of the moment. The machine fires the right model the instant you name it, with no warm-up and no fumbling.
AI doesn’t have the same limited working memory. You can hold only about three or four models in your head at once. The machine holds 750,000 words in its working memory. Ask it to run 10 in sequence on one decision, then 10 more, and it does not get tired, and it does not drop the thread.
AI doesn’t have emotions. You cannot apply a model when the feeling is running the show. The machine has none. It applies sunk cost to the project you cannot bear to kill and says no to the request you cannot bring yourself to decline, because it is not the one who has to feel it. This helps reveal thoughts your emotions would hide.
Next, I decided to move to a higher leverage approach...
A Few Paragraphs That Upgrade Every Answer By Default
I stopped retyping the instruction onto every question, and I wrote a system prompt.
A system prompt is just the block of instructions a model reads before it reads anything you type. Both Claude and ChatGPT allow you to create one in the settings pages of each. So instead of adding that one instruction onto the end of every question by hand, I wrote it in the settings pages once. Now it runs on every conversation, whether I remember to ask for it or not. I no longer have to decide, question by question, whether to think this way. It happens on its own.
Here is what it does. You ask it anything, and instead of one answer, you get four:
#1. plain answer
This is the response the model would have given on its own, untouched, before any of this fired. I keep it on purpose. It is the control, and the whole point is to see what the rest of it adds.
#2. It shares my cognitive signature
It reads my own thinking back to me based on how I framed the question. It shows me the following thinking I used:
Frame(s)
Paradigm(s)
Mental Model(s)
Then, it shares frames, paradigms, and mental models that were in my blindspot, and it reveals what they would’ve shown me if I had asked the question with these.
All in all, this shows me the assumptions I am working from.
#3. It shows its work (reasoning trace)
It names the sequence it is about to run and walks through it a step at a time: this move, through this lens, and here is what that step turned up. I get to watch the thinking happen instead of being handed a verdict, which is the only way I have ever actually learned a new move.
#4. It provides a model-enhanced answer
This is the one that sequence produces. Not a cleaner version of the first answer. A different kind of answer.
The power of this system prompt is that it provides me with better answers, and it helps me improve my own thinking.
Watch what that looked like on a sample question
The easiest way to show you is to walk through a real one:
How can I 100x the speed and quality of my own learning?
You can get the link to the full answers in the appendix. Below is a summary of each response:
Plain answer: The plain answer came back with a genuinely good list. Six solid pieces of advice, ranked from most to least useful. To summarize them:
Compress your feedback loops
Predict before you read
Hunt for disconfirming evidence
Teach in public
Diversify your inputs
Use AI as a sparring partner instead of an oracle
If I’d stopped there, I’d have been happy. It’s the kind of answer most people would be thrilled to get.
Reasoning Trace: Then I watched it think. This is the part worth slowing down for, because you can see the actual mental models it reaches for.
First, it used first principles: stripping a question down to what’s actually true at the bottom, underneath the usual assumptions. It pulled “learning” apart into the separate jobs hiding inside it:
Information intake
Pattern recognition
Conceptual integration
Skill acquisition (procedural)
Identity/perspective transformation (developmental)
Embodied integration (somatic)
Retrieval and application
Calibration / error correction
Then it ran those through the 80/20 rule: the idea that a small handful of things produce most of the results, and asked which single one would matter more than all the rest combined. The answer caught my attention: not taking in information, the thing I spend most of my time on, but how fast and how often you catch yourself being wrong and fix it.
Next, it questioned the hidden assumptions in my question. (Philosophers have a name for this: Cartesian doubt, deliberately distrusting everything you’ve been taking for granted.):
That “learning” is a single variable.
That faster is better.
That the rate-limiter is me.
Then it used analogical reasoning: borrowing a solution from a completely different field that’s already cracked the same kind of problem. Who else learns incredibly fast, and what do they actually do? It focused on:
Chess players who train against a computer that grades every move the instant they make it.
Surgeons who walk back through exactly what went wrong after a patient dies.
Special-forces teams that pick apart every mission the moment it ends.
High-frequency traders who deploy tight learning loops.
Bayesian forecasters that track every prediction and calibrate over time.
And it stopped itself on the one thing they all share: none of them got faster by taking in more. Every single one got faster by shrinking the gap between making a guess and finding out whether the guess was wrong.
It kept going, a Dunning-Kruger check, an adult-development read, a probe at the edges of the question, before it turned on itself.
Finally it tried to prove itself wrong: the move scientists call falsification, and the fastest way to find the holes in your own thinking. If all of this were right, it reasoned, then for someone like me, dropping things should help more than adding them. I heard that as a quiet accusation: I’d been making a fancier version of a beginner’s mistake, piling up better tools, which feels exactly like progress, instead of changing the one thing that moves the needle.
Enhanced Answer: Then came the answer all that thinking built, and it looked nothing like the list I started with.
It didn’t give me a better list at all. It told me, gently, that I’d asked the wrong question. You can’t make your learning a hundred times faster, it said, because “learning” is too small a thing to multiply like that in the first place. The thing that actually multiplies is the whole setup your learning happens inside.
Fix the setup, and faster learning just falls out of it. Then it gave me a metaphor that drove the point home: trying to learn a hundred times faster is like trying to breathe a hundred times faster. You can’t, and you wouldn’t want to. What you do instead is change the air, build bigger lungs, and let your body quietly start pulling more out of every single breath.
Then it told me three things it recommended changing along with personalized advice on how to make the changes:
Change how fast I find out when I’m wrong.
Use my body to turn thinking into instinct.
Let go of my own ideas, including the ones I’m proudest of, the ones I’ve built my work on.
Bottom line:
Two answers, side by side on the same screen:
The sensible list I started with.
The complete rethinking of the question I was even asking.
It was the same question and the same AI, but mental models completely changed the response.
For Months, I Didn’t Believe My Own Results, But Then The Feedback Started Pouring In
It almost felt too good to be true. I wondered if I was just imagining the improvement. So, I shared an upgraded version of the prompt that tapped into mental models in a more organized way in an article: This “AI Command Language” Upgrades Claude to Opus 5.6.
After publishing the article, the feedback started trickling in. I got more people sharing the impact than almost any other post I’ve shared.
More so, when I talked to people one-on-one, multiple high-level entrepreneurs and executives said they use the skill I created every day and found it extremely helpful (and even life-changing). Here are a few samples:
Joe Stolte, 5-Time Founder With Three Exits
Sean Cushing, Co-Founder of Cantina Creative (clients: Avatar, Atlas, Aquaman, The Marvels, Black Adam, Thor, Captain Marvel, Avengers, The Hunger Games, Guardians Of The Galaxy)
Abhinav Sarapure, Lead Business Intelligence Engineer, HubSpot
Seeing these results, my curiosity was piqued. I wondered...
What are the limits of how smart AI can be with the right “thinking workflow / cognitive harness”?
That’s when I began a new line of research...
The Fascinating Research Behind Cognitive Harnesses
That’s when I saw a lineage of people using a few sentences of scaffolding to help models reason better, and the results have been pretty astounding:
Simply adding the words “let’s think step by step” lifted an older model’s score on a set of math word problems from 17.7% to 78.7%.
Letting the model try several paths and back out of the dead ends instead of marching down one line took GPT-4 on the Game of 24 puzzle from 4% to 74%.
Handing it a small kit of named reasoning moves to run on demand lifted a standard model from 32% to 53% on a hard math benchmark, past a specialized reasoning model that costs far more.
Giving a model tools instead of more size let a 6.7-billion-parameter model match models roughly 25 times larger once it could reach for a calculator and a search engine.
Microsoft Research published a method, SkillOpt, that improves an AI by training the instruction document you hand it rather than the model itself. Freeze the model, optimize the words. It was best or tied-best on all 52 of its tests.
More recently, companies have been moving beyond adding a few words and building entire thinking cognitive workflows on top of AI. The first to really capture my attention is a company called Poetiq. They call what they built a “cognitive harness,” and the term fits the whole category, so it’s the one I’ll use from here.
Here’s what a cognitive harness did for Poetiq.
Poetiq, a six-person team of ex-DeepMind researchers, built a cognitive harness that beat the top AI models on one of the hardest benchmarks in the field. They took the standard Gemini 3 Pro model (not Google’s far more expensive Deep Think reasoning mode) and wrapped it in their cognitive harness. Then they ran it against Deep Think on the very well-respected ARC-AGI-2 benchmark. This benchmark is a set of abstract visual puzzles designed to be simple for a person and punishing for a machine. The average human scores around 66%, while the top models were stuck in the single digits when it launched. It rewards figuring out a brand-new pattern on the spot, not recalling something from training.
Deep Think had just set the top of the leaderboard at 45%. Then, Poetiq scored 54%, verified by the ARC Prize team, at less than half the cost per problem ($31 versus $77).
Google pours billions into training that model and employs some of the most decorated researchers in the field. Poetiq is six people. No custom training data. No hundred-million-dollar compute budget. They didn’t build a better brain. They built better thinking on top of a brain.
That last line is worth slowing down on. A cognitive harness is a layer of structured thinking that sits on top of an AI model and organizes how it reasons: which cognitive moves to make, in what order, and how to check its own work. The model underneath never changes. It’s the same Gemini or Claude everyone else types into. The harness is the judgment you wrap around it, so the machine works the problem the way an expert would.
Source: Poetiq
The interview clip with co-founder and CEO Ian Fischer sealed the deal for me:
Source: Y Combinator
Even then, I hesitated to build a cognitive harness, because so many of its components are copyable that could get absorbed into the next model. But the deeper I went, the more convinced I became that the part that compounds, the opinionated judgment you encode, is the part the labs can never absorb. Fischer explained it best.
The conventional approach to making AI smarter in your domain is fine-tuning: you collect a massive dataset, spend months and millions training the model on it, and end up with something that performs a bit better on your specific problem. Then the next model comes out. And it’s better than your fine-tuned version out of the box. As Fischer put it:
“You’re going to spend a lot of money on that fine-tuning. The compute is so expensive. And then at the end of it, you have something that works better than the thing that you fine-tuned on top of, but by then, a new model’s come out, and it’s better than the thing that you fine-tuned.”
—Ian Fischer, co-founder and CEO of Poetiq, Y Combinator interview, 2026
As a result, the new models lit your fine-tuning investment on fire.
Poetiq took the opposite approach. Instead of modifying the model itself, they built on top of it. They describe the harness as “stilts”: the frontier models aren’t competitors; they’re what you stand on. And the key insight is what happens when the next model ships: that same harness slots right on top of the new model and immediately performs even better. No retraining. No new dataset. The harness compounds.
This is the asymmetry that matters:
Fine-tuning is a depreciating asset. The moment the next model drops, your investment resets to zero.
A cognitive harness is a compounding asset. Every model release makes it stronger.
And that asymmetry is the part the labs can never absorb. The labs absorb techniques, because techniques are identical for every user. A harness is not a technique. It is your unique judgment written down: which thinking moves fire, at which step, on which problems. Consensus can be pretrained into a model. Your deviation from it has no shared signal to train on, because it is yours alone.
The surprising part is what building one doesn’t require: a lab, a research team, a compute budget. What it requires is knowing how to organize thinking in a specific domain. Which is exactly what most experts have spent their careers learning to do.
So, I wondered how I could go beyond a mental model system prompt. Then, I came across a fascinating case study.
A Famous Tech CEO Took AI Mental Models Further Than I Imagined
A single sentence upgraded one answer. A system prompt upgraded every answer.
But both still work one answer at a time.
It became clear to me that the next leap was to wire different models to different stages of a whole process, the right specialist at the right step. By pointing this at the workflow that creates articles like this one, I saw an opportunity to drastically increase the quality and quantity of my writing.
But I had never done something like this before, so I wasn’t sure how to think about doing it exactly.
That’s when I found out about a tool called gstack created by Garry Tan.
Tan is the president and CEO of Y Combinator, the largest tech accelerator in the world. It’s behind companies like Airbnb, Stripe, Reddit, OpenAI, DoorDash, Dropbox, and Coinbase. (Of those, OpenAI is the only one that didn’t come up through a YC startup batch; it launched as the nonprofit “YC Research” while Sam Altman was running YC.) He has been building products for 20 years as a tech entrepreneur, and he is shipping more than he ever has. In one recent 60-day stretch, by his own account, he created three production services and 40-plus features, part-time, while running YC. By his own logical-lines-of-code accounting, he is now shipping at roughly 810 times his 2013 daily pace, and his year-to-date output for 2026 alone already exceeds his entire 2013 total by about 240 times, which is still absurd.
One of those tools was gstack, and it has crossed 120,000 GitHub stars, making it the 77th most-starred GitHub project of all time. Open the repository expecting clever engineering and you find almost none. It is nearly all plain text.
What is inside is not a prompt. It is a company. gstack turns one AI into a team of 23 specialists and eight power tools that runs every project through the same seven-step sprint, each step handing its work to the next:
Think. Pressure-test the idea and decide what is worth building.
Plan. Turn it into an architecture and a scope.
Build. Write the code.
Review. Hunt for bugs and weak spots.
Test. Run it for real and probe it for security holes.
Ship. Version, document, and release.
Reflect. Log what was learned for next time.
The engine of each step is not code. It is mental models. At every step, the specialist running it fires a specific set of named thinking tools, each wired to the moment it helps. In planning alone, three different reviewers each fire a different shelf. Ask the CEO reviewer to pressure-test a plan and it runs, in order, Jeff Bezos’s test of whether a choice is a reversible “two-way door” or an irreversible “one-way door,” Charlie Munger’s inversion (”what would make this fail?”), Steve Jobs’s focus as subtraction (he cut Apple from 350 products to 10), and Sam Altman’s question of where the real leverage is. Hand the same plan to the engineering reviewer and a different shelf fires: Dan McKinley’s “every company gets about three innovation tokens,” Martin Fowler’s “refactor, don’t rewrite,” Fred Brooks on separating real complexity from the kind you invented. Hand it to the design reviewer and you get Dieter Rams, Don Norman, and Steve Krug’s “don’t make me think.” Counted up, that is dozens of named mental models, borrowed from dozens of different operators, each assigned to the moment where it earns its keep.
This is the difference between using a model and building a system. A system prompt fires your models on one conversation. gstack fires the right models at the right step across an entire pipeline, every time, with no one having to remember to. None of it made the underlying model smarter. It is the same model everyone else types into. What Tan did was write his whole judgment down as named heuristics wired to a sequence, so the right lens fires on every task instead of only on the days he remembers to reach for it. He gave the machine his mind as a system, and it now runs that mind at a width, and a consistency, he never could.
Tan did this for code. I did it for the one thing I have given my life to: thinking and writing with mental models. Earlier I confessed that even with a head start, I could never run the latticework in my own head. This is what I did instead of giving up on it: I built a system to run it for me.
Everything I have learned and coined about how to think and write, I have been writing down as a set of skills an AI runs. Dozens of them now. One carries my writing voice. One runs the analytical model I use to read the news. Others hunt for quotable lines, pressure-test a draft sentence by sentence, and sharpen every subheader. Each is a piece of my judgment in my core specialty, wired to the moment it helps, so my best thinking is present by default instead of only when I remember to reach for it. The article you are reading went through that system. It is not a description of the method. It is the method’s output.
What Tan and I both stumbled into is bigger than code or writing. It is the clearest case yet of a kind of thinking that was never possible before, and the cleanest way to see it comes from the most influential psychologist of the last 50 years...
Beyond Nobel Laureate Daniel Kahneman’s System 1 And System 2 Thinking. Meet System 3.
Daniel Kahneman is famous for popularizing System 1 and System 2 thinking in his New York Times bestselling book, Thinking Fast And Slow.
In it, he distinguishes between two types of thinking:
System 1 relies on our associative memory. It uses connections between earlier seen words, images, actions, and emotions to form a conclusion and make quick decisions. In certain situations, System 1 thinking can even be more accurate than System 2 thinking. In other situations, it makes silly errors.
System 2 thinking is the slow, deliberate, and analytical mode of human cognition. It requires conscious mental effort, focus, and energy, making it responsible for complex problem-solving, rational decision-making, and critical thinking.
Using Charlie Munger’s system with AI falls into a third category: System 3 thinking. System 3 is using AI not to replace your thinking, but to augment it: holding more than any one mind can hold, while structuring the thinking process according to your own expertise. You can think of System 3 thinking as ultra-thinking where you:
Look at something from many perspectives and paradigms.
Identify and challenge assumptions.
Use multiple mental models in sequence.
Attempt to cancel out every blindspot and bias.
System 3 has the speed of System 1 and the intelligence of System 2. And, as AI gets better, the gap between System 3 and Systems 1 and 2 will increase.
For all of human history, System 3 was a fantasy. No mind was ever deep enough, fast enough, or wide enough to hold it. That is the real reason almost no one could run Munger’s latticework. Not for lack of wanting. For lack of hardware.
The implications of this are profound…
Implication #1: Cognitive Harnesses Have The Potential To Usher In A New Era Of Democratized And Amplified Intelligence
Remember the question that guided my research a decade ago: the horse-drawn plough reduced the importance of physical strength in the world, so what is the plough for the mind?
I can finally give you my answer. The cognitive harness is that plough.
The reason that question mattered to me went beyond my own career. If a tool could augment intelligence, then the advantages of intelligence could finally be spread.
And the plough analogy is more precise than it first appears. The plough didn’t make a single farmer stronger. It made strength less decisive. Before it, the land a family could work was governed by muscle. After it, an ordinary farmer behind a plough out-produced the strongest man in the village working without one. The tool didn’t upgrade people. It downgraded the trait.
A cognitive harness does the same thing to raw intelligence. It doesn’t raise your IQ. It makes raw IQ less decisive, because the heavy cognitive lifting — storing thousands of models, retrieving the right ones at the right moment, firing them in sequence, checking the work — moves out of your biology and into the system. At the same time, it also helps make you smarter.
To feel why that’s a democratizing force, look at who amplified intelligence has belonged to until now.
To run even a partial latticework, you had to win several lotteries at once:
The genetic lottery. A memory and a processing speed most people are not born with.
The time lottery. A job where thinking is the work, so you can spend most of the day reading, the way Buffett and Munger famously did.
The money lottery. The wealth that buys that kind of day in the first place.
The temperament lottery. Decades of patience for a payoff that compounds slowly and invisibly.
History shows what happens when a scarce ability like that finally gets packaged into a tool:
Before the printing press, the accumulated knowledge of humanity belonged to whoever could afford hand-copied manuscripts, which is to say, almost no one. Then the price of a book collapsed, and within a few generations, so did the monopoly on knowledge.
“Computer” was a job title before it was a machine. Complex calculation belonged to trained specialists, and rooms full of human computers powered observatories and space programs. Then the machine version arrived, calculation became effectively free, and today nobody brags about being able to do long division.
Each time, the pattern repeats. The scarce ability doesn’t just spread. It stops being the thing that separates people, and the frontier moves up to whatever the tool still can’t do.
Now recall the two charts near the beginning of this article: a century of rising IQ, then the recent reversal. Every force commonly credited for that century of gains — better nutrition, more schooling, smaller families, more abstract work — operates on the scale of generations.
A cognitive harness is the first intelligence lever in history that operates on the scale of an afternoon. You install it once, and the very next question you ask is processed through more models than any unaided mind has ever run.
Think about what that makes possible for people who were locked out of elite thinking until now:
The first-generation college student gets the decision checklist that used to require a mentor at a hedge fund.
The solo founder in a small town gets the strategy pressure-test that used to require a consulting retainer.
The mid-career manager gets a Munger-grade review fired on every major call, not just the ones she remembers to slow down for.
The family facing a hard decision gets their reasoning stress-tested from ten angles before they walk into the room where it gets made.
To be clear, democratized does not mean flattened. Amplifiers never abolish advantage; they relocate it. The advantage stops living in what you were born with and starts living in what you build on top of the amplifier. More on that in a moment.
But make no mistake about the direction of the change. The floor is rising faster than at any point in the history of thinking. Anyone willing to learn the system can now run a wider latticework than Munger ever did — because Munger had to run his only on three pounds of tissue, and you don’t.
Implication #2: The Only Ceiling Left Is the One You Set Yourself
Mine was handed to me at eight years old, one point short of a program, and I unconsciously spent the next 30 years trying to outwork it.
By the test’s own logic, I was never “not smart.” One point short of the gifted line puts you near the top. The number was never what held me back. What I decided it meant did.
For most of history, most verdicts like that came with an alibi. You could point to your genes, the “hardware” you were issued, and say the ceiling was real and external. The latticework stayed out of reach for lack of hardware, and that was physics, not failure.
The alibi is gone. The limit that made the latticework impossible has been lifted for almost everyone at once. The ceiling that remains is no longer the one reality assigned you. It is the one you assigned yourself. You can think far past what you were told, and you can no longer blame the number if you don’t.
We are drawing a new line right now, not by genes this time, but by who learns to run these systems and who does not. The person fluent in System 3 thinking will out-decide, out-create, and out-earn the person who takes the first answer the machine gives, and the gap will look like the old one between the gifted and the rest.
With all that said, I want to officially announce and share the second version of the mental model cognitive harness…
Announcing Version 2 Of My Mental Model Operating System
Earlier this year, I released a skill called Mental Model Recipes. It organized more than 100 mental models into 40 thinking “recipes,” sequences of structured moves that gave Claude answers with depth and nuance that the base model couldn’t reach on its own. It became the most popular AI skill I’ve ever shared.
This is version 2 created in collaboration with Bonnie Johnston.
Here’s what’s new…
When you ask a question, it doesn’t just answer it. It runs your question through a six-step process before it ever gives you a response. Each step sharpens its thinking. Let me walk you through what happens under the hood.
Step 1: It diagnoses what you’re actually stuck on.
Not what you asked, but what’s underneath the asking. The skill classifies your situation into one of 10 kinds of stuckness: you can’t choose, you can’t create, you can’t diagnose, you can’t predict, you can’t persuade, you can’t understand, you can’t execute, you can’t align, you can’t grow, or you can’t change the system.
This matters because the kind of thinking tools you need depends entirely on what kind of stuck you are. A decision problem and a creativity problem need completely different mental models.
Step 2: It identifies the worldview you’re operating in.
This is the water you’re swimming in: the paradigm, the assumptions, the mental map of reality you’re using without realizing it. Your worldview shapes what you notice and what you miss, and it’s almost invisible from the inside.
Step 3: It fetches the opposition.
Here’s why this step exists. There’s a catch to having one AI look at your question from many perspectives: all of those perspectives come from the same mind, and that mind wants to agree with you. Left to its defaults, AI bends every lens toward the way you framed the question. You can get what feels like a unanimous panel and is actually one voice wearing 300 costumes. And that false unanimity is most dangerous at the exact moment you need dissent most: right before a big decision.
So we built the disagreement in. From 338 rival pairings, the skill finds a worldview that sits in tension with yours and asks: what does this perspective see that yours doesn’t? This is the step that didn’t exist in v1, and it’s the one that surprises people the most. You’re not getting a second opinion from the same angle. You’re getting a genuinely different lens, one that reveals what your own worldview hides. It breaks the echo chamber you didn’t know you were in.
It also gives you a reading rule for every answer the Stack produces: trust the divergence, audit the unanimity. Where the lenses disagree, that fork is the signal. Where they all agree with you a little too smoothly, that’s the moment to slow down.
Step 4: It de-biases your question, not just the answer.
The previous skill de-biased the answer. This one starts a step earlier. It looks at the cognitive biases and logical fallacies implied in the way you framed the question, and then it asks: if you didn’t have those biases operating, what other questions might you have asked? It answers those questions too, and lets those answers shape the response you get to your actual question.
Step 5: It runs the matching thinking recipe.
Based on everything it now knows (your stuckness type, your worldview, the rival perspective, the de-biased question), it selects one of 61 recipes: a specific, ordered sequence of thinking moves built from the full library of 2,375 mental models. This isn’t “here are some mental models that might be relevant.” It’s a structured sequence, chosen for this specific situation, executed in order.
And yes, do the math on that library. The title of this article promised your AI could run 300 mental models. I undersold it. The full build chooses from 2,375.
Step 6: It checks its own work.
It tracks each operation it performed (each model applied, each paradigm surfaced, each bias corrected) and verifies that each one actually changed the answer and added something that mattered. If a step can’t beat the plain answer, the skill is built to say so. Most AI tools never tell you when they’re not helping. This one does.
What you get back is the four-part answer you saw earlier in this article, plus a fifth part:
A “natural” answer. The one you would’ve gotten from the AI without the Mental Model OS, so you have a baseline for evaluating the enhanced answer.
A mirror of your own thinking patterns. Your cognitive signature.
The full reasoning trace. Showing what was selected and rejected, and why.
The recipe-driven answer. Not a cleaner version of the first answer. A different kind of answer.
A concrete next move you can act on.
One note on when to reach for it. The OS is built for thinking-type questions: decisions, diagnoses, strategy calls, the moments you’re genuinely stuck. For a quick factual lookup it’s overkill, which is why the install gives you a one-word off switch: start your message with “Answer:” and it stands down. And it isn’t fast. You’re trading a one-line reply for a five-part one, on purpose. When a question is worth real thought, that’s the trade you want. When it isn’t, don’t use it.
From 100 Models To 2,375. And That’s Not Even The Big Upgrade.
If you’ve been using Mental Model Recipes, you already know the skill works. Version 2 is what happens when you take that foundation and ask: what’s still missing?
The numbers tell one story. But the three upgrades you’ll actually feel are the ones that don’t have a v1 equivalent at all:
It argues against you. V1 gave you better mental models for your question. V2 goes looking for a worldview that disagrees with yours and asks what it sees that you don’t.
It fixes the question first. V1 corrected for bias in the answer. V2 starts a step earlier, stripping the bias out of the way you asked before it responds.
It grades its own work. V1 had no self-verification. V2 confirms each step actually improved the answer, and tells you when one didn’t.
The short version: V1 gave you the right mental models for a question. V2 first figures out what you’re stuck on, shows you the worldview you didn’t know you were in, drags in the opposition, strips the bias out of the question itself, then checks its own work.
The Test Most AI Skills Would Fail
There are a lot of AI skills floating around the internet right now. The problem is:
You have no idea what thought was put into them
You have no idea how specifically they’re designed to fit the creator’s situation or personal workflow
You can’t always tell if the results are better than what the AI would naturally do with a simple prompt
They don’t check their work
They just run a process
We didn’t want to ship something like that.
So we built a testing rubric. We gave it a batch of questions, some we generated ourselves, some were real questions we actually needed answers to, and scored the results. The rubric tested each operation the skill performs in sequence:
Did this mental model actually change the answer?
Did this paradigm surface something the previous step missed?
Did the bias correction add something that mattered, or was it decoration?
The rubric goes deep. For each step in the process, it evaluates whether that specific operation contributed meaningfully to the final answer. Not “was the answer better,” but “was it better because of this step.” That’s a harder question, and it’s the one that matters. If a step doesn’t earn its place, the skill should know it, and so should you.
We also built this self-checking into the skill itself. It doesn’t just run the rubric during testing; it runs a version of that verification every time it answers a question. That’s an unusual feature for any AI tool: honesty about its own limits.
Two Editions
The OS comes in two editions, built for two different ways of working with AI.
The Chat edition is the curated build: 561 mental models, the full paradigm and rival-pairing system, and the six-step router. You upload it like any Claude skill, and it works in Claude Chat. This is the version to start with if you primarily work in the chat window.
The Code edition is the full build: all 2,375 mental models, the complete data library, a Python engine, and the ability to split work across multiple agents for deeper analysis. It runs in Claude Code, or in Claude Co-work if you’re newer to the agentic side. This is the version for people who are already working in the agentic paradigm or want to be.
The simple choice: if you live in the chat window, start with Chat. If you’re building in Claude Code, go with Code. Or install both and never miss an opportunity to upgrade your thinking again.
The download links and step-by-step install instructions for both editions, including the auto-trigger setup so the Stack fires on its own, are in the paid section at the bottom of this post.
Paid members of this newsletter get version 2 of my mental model operating system, the most popular AI skill I’ve ever shared. You can access it at the bottom of this post after the appendix.
APPENDIX
#1. The Mental Model Studies
Study #1: Experts think about what mental models are relevant to a problem before trying to solve it. Amateurs don’t.
The first thing the research shows is that expertise itself is largely model-based.
In 1981, Michelene Chi, Paul Feltovich, and Robert Glaser handed physics problems to both novices and experts, then asked each group to sort them.
The novices sorted by surface, matching inclined planes with inclined planes, pulleys with pulleys.
The experts ignored the surface and sorted by the deep principle that solved the problem: conservation of energy in one pile, Newton’s second law in another. They weren’t seeing a harder version of the same picture. They were seeing through models the novices did not yet have.
Study #2: Experts with more diverse mental models outperform experts with few models from one field.
The bigger claim, the one this whole piece rests on, is that carrying many models from many fields beats carrying one. The largest test of this principle ran for two decades. The psychologist Philip Tetlock tracked 284 experts making roughly 28,000 forecasts about politics and economics and scored every prediction they made
against what actually happened.
Most barely beat chance. One group did consistently better.
The generalists did better. Referred to as “foxes,” they pulled from many disciplines and drew on several models at once.
The specialists did worse. Labeled “hedgehogs,” they ran everything through one big theory.
The hedgehogs didn’t just fail to outperform. They were worse, and more overconfident, especially on the long-range calls inside their own specialty. One powerful model, applied to everything, was a liability. The cure was range.
Study #3: The most creative experts have the widest range of interests.
Tetlock measured the impact of range on judgment. But physiologist Robert Root-Bernstein measured the impact of range on creativity.
Root-Bernstein spent decades comparing eminent scientists with ordinary ones, and the gap was not raw ability. It was range. Two studies particularly stand out:
Against a baseline of typical scientists, Nobel laureates in the sciences turned out to be far more likely to pursue the arts: roughly twice as likely to play an instrument, around five times as likely to work a craft, and more than 20 times as likely to act or dance, with visual art and creative writing falling in between.
In a 25-year study that followed 40 young scientists, the ones who went on to win Nobel Prizes or election to the National Academy of Sciences had the widest range of outside pursuits, and could explain how those pursuits fed their science. The least successful had the fewest, and treated outside pursuits as distractions.
Range matters because a model carried in from one field is often the missing key in another.
Einstein said it plainly:
“If I were not a physicist, I would probably be a musician. I often think in music. I live my daydreams in music. I see my life in terms of music.”
He claimed he rarely thought in words at all, only in images and feelings that he translated into equations afterward. The theory that broke open physics was built, in part, with tools borrowed from music.
Study #4: Your range of outside interests predicts success better than your IQ.
Range does not just help at the margin. It can outweigh the very things we are taught to measure. In an 18-year follow-up study, the psychologist Roberta Milgram found that a person’s intellectually stimulating, intensive outside interests predicted their career success better than IQ, grades, standardized test scores, or all of them combined. The hobbies pursued out of genuine curiosity, the ones that look like detours, forecast achievement more reliably than the scores we spend an entire education chasing.
Study #5: Storage of mental models is not enough. Retrieval is key.
The last three studies make range sound like a free lunch. It is not. Having the models is not the same as using them.
In a classic experiment, Mary Gick and Keith Holyoak gave people a problem almost no one solves cold: destroy a tumor with rays without destroying the healthy tissue around it. Fewer than 1 in 10 cracked it. But people who had first read an unrelated story, about a general taking a fortress by sending small forces down many roads at once, and who were nudged to use it, solved the medical problem at rates of 75% to over 90%. One model, carried in from war and mapped onto medicine, was the difference between stuck and solved. The same study delivers the most important caveat in this article. When people got the fortress story but were not told to use it, only about 20% made the leap on their own. Holding the right model in your head is not the same as reaching for it at the right moment.
Said differently, the science does not say that cramming your head with models from every discipline automatically makes you smarter. Mostly it says the opposite: the hoped-for “far transfer,” where training one skill lifts unrelated ones, barely shows up once studies are run with proper controls. The polymaths in those studies were not passive beneficiaries of their range. They actively reached into one field, pulled out a model, and mapped it onto a problem in another, which is the move that pays.
The bottleneck was never storage. It was retrieval.
Study #6: The basis of expertise is the number of stored mental models.
The pattern holds across fields.
Chess masters do not have better memories than the rest of us. Shown a real game position for a few seconds, a master rebuilds most of the board while a novice gets only a handful of pieces. Scatter those same pieces at random and the master’s edge nearly vanishes, because what the master stored was never the pieces. It was the patterns, the models of how pieces relate.
More recently, researcher Daniel Simons replicated the experiment with chess grandmaster Patrick Wolff and posted it on Youtube. Wolff’s attempt in the random position is almost pathetic. In the real position, it is a miracle. Only two pieces are misplaced, and they’re only misplaced by one square.
Also, notice how Wolff describes the feeling when he’s chunking. Rather than describing the process of working with chunks as difficult, mechanical, or deliberate, he describes it as flowing and unconscious. This is a hallmark of mastery.
Study #7: The pattern applies to teams as well
The same logic scales past the individual.
In a widely cited (if contested) modeling result from Lu Hong and Scott Page, a cognitively diverse group, people running different models, can outperform a group of the highest-ability individuals on hard problems, because where one model is blind another can see.
What range does inside one head, diversity does inside a team. A room full of the same brilliant model is still one model, blind in all the same places. A room running different models sees around its own corners. Range beats brilliance, whether it lives in one mind or many.
Sources:
https://docs.google.com/document/d/16MrG2LBoQZHj87eclTQaiXO2Dy7Aa5WDedjBIRfWY_Q/edit?tab=t.0
https://docs.google.com/document/d/1U1OvHDmB3dX8P3LSXR6fJp9kDqYnvn5Tj304lEtqIYU/edit?tab=t.0
#2. IQ Research
An extra hour of sleep has a surprisingly large impact on IQ. For example, the performance gap caused by an hour’s difference in sleep is bigger than the gap between a normal fourth-grader and a normal sixth-grader. Which is another way of saying that a slightly-sleepy sixth-grader will perform in class like a mere fourth-grader. “A loss of one hour of sleep is equivalent to [the loss of] two years of cognitive maturation and development,” pediatric sleep specialist Avi Sadeh told the authors of NurtureShock. Sadeh was the lead author on a 2003 study that monitored swings in the cognitive performance of 77 school-age children. Neuroscientist Dr. Tara Swart cites research showing that disturbances to healthy sleep patterns can cause your IQ to drop by 5 to 8 points, and that missing out on just a single night of sleep can reduce your IQ by one standard deviation, or 15 IQ points.
Regular exercise improves neuroplasticity and memory while slowing cognitive decline. Research suggests that physical exercise may trigger processes facilitating neuroplasticity, enhancing your capacity to respond to new demands. One year of aerobic exercise in a large RCT of seniors grew hippocampal volume and improved spatial memory, with other trials showing slowed gray-matter loss and better brain connectivity over 6 to 12 months (sources).
Creatine helps the brain most when it’s running on empty. In a 2024 trial, a single dose helped sleep-deprived adults hold on to their working memory and processing speed. In well-rested people the effect is small, and regulators haven’t endorsed a general cognitive claim, but under stress, the benefit is real.




















