What a Claude Skill Actually Is, and What It Isn't
A Claude skill is a folder of plain instructions sitting in a GitHub repo, which Claude reads at the start of a task the way you'd hand a new hire a runbook. It doesn't make Claude smarter. It makes Claude stop guessing at things you already wrote down the answer to.
What is a Claude skill, and why does it live in a GitHub repo?
A Claude skill is a folder of markdown files - instructions, examples, house rules - that Claude reads before it starts a task. It sits in a GitHub repo because that’s what a folder of text files wants to be: versioned, diffable, and pull-requestable when it’s wrong. There’s no new model underneath it, no fine-tuning, no separate piece of software running alongside Claude. It’s the same Claude you already have. It just walked in having read the manual first.
That’s the whole trick, and it’s a smaller trick than it sounds. A skill doesn’t teach Claude to reason better. It teaches Claude what you already know and would otherwise have to type out again every single time - which repo owns which kind of ticket, what your test command is, which error message means “check the fixture” rather than “check the code”. You’ve written that knowledge down somewhere already, probably in your own head. A skill is that knowledge, externalised, so it survives the conversation ending.
Which is also exactly where it stops. A skill can’t verify a claim it wasn’t told to verify. It can’t notice that the instructions it’s following are stale. And it will follow a plausible-sounding diagnosis just as confidently as a correct one, because “plausible” and “correct” read identically from the inside. We found that out the expensive way this week.
The bug that wasn’t where the error message pointed
One of our jobs died on Encoding::CompatibilityError: invalid byte sequence in UTF-8. A daily log file had a £ sign saved in the wrong encoding, and the process choked reading it back.
The first fix looked obvious. Reproduce it locally, get in US-ASCII instead of in UTF-8, read that as “ah, it’s a locale problem, the process isn’t picking up the right character set” and ship a fix that exports the locale explicitly. Confident, quick, shipped as a pull request. Job done, or so it looked for about a day.
It was wrong twice over. First, the locale half of that diagnosis had already been fixed four days earlier, in a completely different file, by a completely different piece of work - the “fix” was solving a problem that no longer existed. Second, and more embarrassingly, the original error message never said anything about locale at all. It said UTF-8. Not ASCII. The reproduction produced a different error than the one in front of us, and that difference got waved through instead of questioned, because the new error was in the same neighbourhood and neighbourhood was close enough to feel like confirmation.
The actual fix, once someone went looking properly, was to scrub the byte on the way in, at the one place a file gets read from disk - a few lines, pinned down with a test that writes the exact byte that broke it in the first place, so it can’t silently come back. Nothing to do with locale. Nothing to do with the environment the process runs in. Just: don’t trust a byte you haven’t checked.
The pattern worth stealing
None of that is specific to Claude, or to skills, or to us. It’s the shape of every misdiagnosed bug anyone has ever shipped: an error message that sounds like a category you recognise, a fix for that category that’s sitting right there, and the actual reproduction quietly swapped out for a reproduction of something adjacent. The fix goes out. It’s wrong. Nobody notices immediately, because “wrong but shipped” and “right” look the same in a green CI run.
You have a version of this in your own history. Everyone does - the config change that fixed a symptom nobody asked about, the retry logic added because the timeout “felt” like a network issue. The tell, in hindsight, is always the same: the error you reproduced wasn’t the error you started with, and nobody stopped to check.
That’s the one rule worth carrying forward, and it’s small enough to actually remember: reproduce the exact error in front of you, not an error that resembles it. If your local repro says something different from your production log, that gap is the finding - not a footnote to skip past on the way to the fix you already had in mind.
That rule now lives in the same place the rest of this knowledge does - written down once, in a file, so the next session starts already knowing it instead of relearning it live. Which is the honest answer to what a skill is for. It doesn’t stop you from being confidently wrong. It just makes sure that when you finally work out why you were wrong, you only have to write it down once.