Read time: 4 minutes | Unlock my Claude Skills Bundle
Hey friend, Alex here.
Last November, Andrej Karpathy wrote about 80% of his code by hand. Six weeks later the ratio had flipped: agents write 80%, he edits the rest. The man who coined “vibe coding” now says he’s mostly programming in English.
The same thread carried a warning that got far less attention than the flip. The agents make the mistakes of a sloppy, hasty junior dev: wrong assumptions they never check, confusion they never surface, 1,000 lines of code where 100 would do. His advice was to watch them like a hawk. Nobody has time to watch every line, so the junior-dev mistakes ship straight to production.
Today:
The Karpathy rules turned into a Claude skill that makes your coding agent behave like a senior engineer (one file, paste and install)
Karpathy’s own field report: what changed when agents started writing 80% of his code, and the mistakes that survived
4 resources worth saving, including the viral GitHub repo built from the same Karpathy thread
⚡ THE SUPERPOWER: The Senior Engineer Skill
If you’ve run Claude Code or Cursor for more than a week, you’ve met the junior dev. You ask for a bug fix and get half the file rewritten. You ask it to add one validation check, and it builds a validation framework with a config file you never requested. You float a bad idea just to think out loud, and it answers “Of course!” and starts implementing.
Karpathy catalogued all of it in that January thread: wrong assumptions run with unchecked, no clarifying questions, no tradeoffs presented, sycophancy, bloated abstractions, and dead code left behind after refactors. His verdict: these errors pass the compiler every time. The failures live in judgment now.
And that points at the fix. Your agent already has the knowledge of a senior engineer. What it’s missing is the judgment, and judgment is a set of habits you can write down. Knowledge comes free with the model. Judgment, you have to install.
Karpathy already ran the field test. Weeks of heavy agent coding on his own projects convinced him the shift is the biggest change to his workflow in two decades of programming, and every failure mode this skill targets came out of those real sessions.
⚙️ THE WORKFLOW
So I went through the thread, complaint by complaint, and turned each one into a rule. Wrong assumptions became a mandatory “here’s what I’m assuming, correct me now” block before any non-trivial change. Sycophancy became an explicit duty to push back. The 1,000-line problem became a simplicity check that the agent runs on itself before finishing. Twelve failure modes in, twelve guardrails out.
It lives in one file. If you use Claude Code, drop it into the CLAUDE.md at your project root, and it loads with every session. If you work in a Claude project on the web, add it to the project knowledge as senior-engineer.md instead. Same rules, both doors.
# Senior Engineer Skill
You are a senior software engineer working alongside a human who
reviews everything in an IDE. You are the hands; the human is the
architect. Move fast, but never faster than they can verify.
## Before any non-trivial change
State your assumptions explicitly:
ASSUMPTIONS I'M MAKING:
1. [assumption]
2. [assumption]
→ Correct me now or I'll proceed with these.
Never fill in ambiguous requirements without saying so.
## When confused
Stop. Do not guess. Name the specific confusion, present the
tradeoff or ask the question, and wait. "I see X in file A but Y
in file B. Which takes precedence?" beats a coin flip every time.
## Push back when warranted
You are not a yes-machine. If the human's approach has a clear
problem: name the issue, give the concrete downside, propose an
alternative, then accept their call. "Of course!" followed by a
bad implementation helps no one.
## Keep it simple
Before finishing, ask: can this be fewer lines? Are these
abstractions earning their complexity? If you wrote 1,000 lines
where 100 would do, you failed. Prefer the boring solution.
## Surgical scope
Touch only what you're asked to touch. Don't remove comments you
don't understand, don't refactor adjacent systems, don't delete
code that merely looks unused. After a refactor, list any dead
code and ask before removing it.
## For multi-step tasks
Emit a short plan first (step + why), then execute unless
redirected. For non-trivial logic, write the test that defines
success before the implementation. Naive-but-correct first,
optimize second.
## After any change, summarize
CHANGES MADE / THINGS I DIDN'T TOUCH / POTENTIAL CONCERNS.
## Communication
Be direct about problems. Quantify ("adds ~200ms latency," never
"might be slower"). When stuck, say so and show what you tried.
Don't hide uncertainty behind confident language.
Once it’s installed:
Start a real task, then read before you run. The first thing back should be an assumptions block or a plan, never code. That pause is the skill working.
Answer the assumptions. Even a one-word “correct” or “no, use Postgres” steers the whole implementation. Those 10 seconds save you more rework than anything else in the session.
Test the spine once. Suggest something you know is a bad idea. A stock agent agrees. This one should name the downside and offer an alternative. If it caves, the file didn’t load.
Make it yours. Add your stack’s conventions and forbidden patterns to the file. The rules above are the floor, your context is the ceiling.
The proof this week comes from Karpathy’s field notes rather than my own lab.
He ran the shift on real projects: roughly 80% manual coding in November, 80% agent coding by December, with the failure modes above catalogued from weeks of daily sessions.
Two of his observations map straight onto this file. Challenged on bloat, agents cut 1,000-line implementations down to 100 on request, which is why the simplicity check runs before finishing instead of after you complain. And their persistence is real, they’ll grind on a problem for 30 minutes without getting demoralized, which is why the file forces them to confirm the goal first. Persistence on the wrong problem is just faster failure.
Fair warning: the rules add overhead on trivial work. Ask for a one-line fix and you’ll sit through an assumptions list you didn’t need, so keep a plain chat around for small stuff. And on long sessions the guardrails fade as the context fills up; when the agent starts agreeing with everything again, remind it the skill exists.
The condensed version above keeps every rule from my full system prompt. If you want the complete original with the XML structure and all 12 named failure modes, it’s pinned in the source post linked below.
💬 PROMPT OF THE DAY
Read the email I'm about to send. Tell me what the recipient will
actually DO after reading it. If the answer is "nothing" or "reply
with a question," rewrite it so the next action is obvious.
Why it works: it grades the email on the recipient’s next action instead of on tone, which is the thing most drafts fail at.
📚 USEFUL RESOURCES
🧵 Karpathy’s original thread (the 80/20 flip post) → the primary source for today’s skill, worth reading in full for the “slopacolypse” prediction alone (free: x.com/karpathy/status/2015883857489522876)
📦 forrestchang/andrej-karpathy-skills (viral CLAUDE.md repo) → a four-principle distillation of the same thread, good to compare against today’s version (free: github.com/forrestchang/andrej-karpathy-skills)
📖 The 80% Problem in Agentic Coding (Addy Osmani essay) → the clearest writeup of what changes when agents write most of your code (free: addyo.substack.com/p/the-80-problem-in-agentic-coding)
🛠️ Claude Code memory docs (how CLAUDE.md files load) → five minutes here explains why today’s install method works (free: code.claude.com/docs/en/memory)
What should I turn into a skill next: the NotebookLM research prompts or the Lead Software Architect prompt? Hit reply, one word is enough.
Know someone who merges agent code without reading it? Forward them this.
And as always, remember: LLMs don’t think, you do.
⚡ Alex Prompter
P.S. Today’s Senior Engineer skill is one file. My Claude Skills Bundle turns Claude into 20+ specialists for marketing and business, each one installing judgment the same way. Grab yours in one click: linktr.ee/alex_prompter



NotebookLM
Notebook