Beer & Servers Don't Mix

The Skills Sprawl Problem Nobody Was Looking For

Or: How 56 Repos Quietly Appeared While We Were Busy Being Impressed With AI

The Sourcegraph MCP query had been running for maybe four seconds when the terminal started filling. Repo names, one after another, in that flat monospaced cascade that always feels faster than it is. I leaned back in my chair on level 6, the Thai iced tea sweating onto the desk next to my keyboard, and watched the count tick past forty. A colleague had asked for the data. I’d told him a week earlier that I thought we had a problem with skills sprawling across the org, and he’d done the thing he always does — not interested in opinion, wanted data. The number stopped at fifty-six.

We had a problem.

This is a post about that problem, and about the unglamorous platform-engineering work it took to turn it into something that looks vaguely like a solution. It’s a post about how an open standard, a couple of MCP clients, and a few thousand enthusiastic engineers can produce a kind of organisational mess you didn’t even have a name for yet — and how the answer, when you find it, looks suspiciously like every internal-platform problem you’ve already solved twice in your career, just with an LLM sitting where the human used to be.

How We Got Fifty-Six Repos Without Trying

Anthropic published the Agent Skills format in late 2025 (here’s the original write-up). By December it was an open standard at agentskills.io. Cursor shipped Skills support in version 2.4 in January 2026. Claude Code had filesystem-based skills from day one. Within a few months, every coding agent in our stack — and probably yours — had agreed on the same primitive: a folder with a SKILL.md in it, YAML frontmatter at the top, prose below, optional resources alongside. Wonderful. The standards war that everyone feared simply… didn't happen.

Which meant the bottleneck moved. The question stopped being “how do I teach my agent” and became “how does anyone else find what I taught my agent.” That’s a much harder question, and it’s the one nobody was asking yet.

Here’s what was actually happening at Agoda in April 2026, in numbers:

  • A few thousand engineers- Roughly 75% monthly active on Claude Code- Roughly 95% monthly active on Cursor- 56 repositories containing skills- Roughly 70% of those skills shipped as raw markdown, the rest as Claude Code plugins- Zero shared way for any of those skills to reach an engineer who didn’t already know they existed The shape of that sprawl is worth dwelling on, because it’s not specific to us. It’s the same shape every shared-library sprawl in your career has had. The platform appears. Early adopters experiment. The successful patterns get copied — repo forked, file pasted, slightly edited, committed. Now there are two versions of the same skill and no signal which one is canonical. Multiply by a few thousand engineers, three coding agents, and a year of organic adoption, and you get fifty-six repos. You got there the same way NPM got its left-pad moment and the same way every Java shop in 2008 ended up with seven slightly different utility libraries: through a series of locally rational decisions in the absence of a centralised path.

The economist Thomas Schelling, who never wrote a line of code in his life, put it better than any engineer I’ve read: “What is rational for each member of a group taken separately may be irrational for them taken together.” Every single engineer who forked a skill repo made the right call. The aggregate was a mess.

We tried Cursor’s native rules-and-prompts sharing first, by the way. It failed for a reason worth naming, because the failure generalises beyond Cursor: it was all-or-nothing. Every engineer in the org received every rule. A backend engineer’s agent context filled up with React component rules. A data engineer’s context filled with marketing CMS guidance. The agent got worse at the actual job, because we were drowning it in irrelevant tokens. And of course the three quarters of the company on Claude Code saw none of it anyway.

The lesson there isn’t “Cursor’s mechanism is broken.” The lesson is structural: any single-IDE skill-sharing mechanism that doesn’t support subsetting will fail at organisational scale. Once you have multiple engineering specialties in one company, you need filtering as a first-class feature, not a workaround. You can’t ask every engineer to manage their own skill list any more than you’d ask them to manage their own package.json by hand for a thousand-package monorepo.

The Shape of the Solution: One MCP, Two Primitives

The reframe that made everything else fall into place was this: skills look like a content problem, but they’re a delivery problem. The content of any individual skill is mostly fine — engineers write good prose when they care, and the people writing skills cared. What we didn’t have was plumbing. We didn’t have versioning, we didn’t have discoverability, we didn’t have permissions, we didn’t have telemetry, we didn’t have tests. We had a pile of SKILL.md files and a vague hope that the right engineer would find the right file at the right moment.

So we built one internal MCP server, and we pointed every IDE at it.

The architecture is embarrassingly simple, which is part of why it works. FastMCP, in Python, scanning a directory of SKILL.md files at startup. Each skill becomes an MCP tool, with the SKILL.md frontmatter as the tool description. An optional resources/ subfolder under each skill becomes MCP resources at skill://{name}/resources/{filename}. The filesystem is the registry. There is no separate JSON config to drift out of sync with reality, because the git diff is the catalog change.

The pitch fits in fifteen words: FastMCP serves skills via tools, and slash commands via prompts. One server, two primitives, every IDE.

That sentence is the whole architecture. Both Claude Code and Cursor are MCP clients. MCP has two native primitives for what we needed: tools for the things an agent invokes by itself, prompts for the things an engineer invokes deliberately (slash commands). We didn’t have to pick a side. We didn’t have to package anything as a Claude Code plugin, which wouldn’t have worked in Cursor anyway. We didn’t have to write a Cursor extension, which wouldn’t have worked in Claude Code. We just ran one HTTP service, and we let MCP do what it was designed to do.

A minimum-viable version is roughly this much code — enough to make the point, not enough to be precious about:

from fastmcp import FastMCP
from pathlib import Path
import yaml
mcp = FastMCP("skills-catalog")
SKILLS_DIR = Path("./skills")
for skill_dir in SKILLS_DIR.iterdir():
    skill_md = skill_dir / "SKILL.md"
    if not skill_md.exists():
        continue
    raw = skill_md.read_text()
    frontmatter, body = raw.split("---", 2)[1:]
    meta = yaml.safe_load(frontmatter)
    @mcp.tool(name=meta["name"], description=meta["description"])
    def _skill(body=body):
        return body
if __name__ == "__main__":
    mcp.run(transport="streamable-http")

That’s it. That’s the v1. If you stripped out the filtering, the telemetry, the tests, and the multi-server topology, this is what’s left, and it would still work. The interesting bits aren’t the code — they’re the four design decisions sitting on top.

Decision one: filesystem is the registry. No central config. A new skill is a new folder; a deleted skill is a deleted folder. There is no truth source for the catalog except the filesystem the server reads at startup. This sounds trivial. It is not. The number of internal platforms that fail because their config drifts out of sync with their content is embarrassingly large, and the fix is always the same: stop trying to maintain two truths.

Decision two: selective serving, not selective installing. The MCP server filters which skills each client sees, based on the team, repo, and project context the client sends. The engineer doesn’t manage a list. The server decides what’s relevant. This is what fixed the all-or-nothing problem we hit with Cursor’s native sharing — filtering moved from the client to the server, where it belongs.

Decision three: multi-server, not mega-server. We didn’t build one giant skills repo. We built several MCP servers using the same pattern. One for engineering coding skills. One for office skills (PowerPoint generation, spreadsheet templates). One for management skills — and yes, that’s a thing, more on it in a second. Each server has its own ownership and its own CI. Backend engineers shouldn’t review PRs to a PowerPoint formatting skill. Business analysts shouldn’t be on the hook for the data platform query skill. This is closer to how UNIX mounts filesystems than how Python imports packages: scoped namespaces that the engineer’s IDE assembles based on context.

Decision four: HTTP traffic is the telemetry. Each skill invocation is one HTTP call. Adoption falls out of the access logs. No client-side instrumentation, no opt-in, no sampling, no analytics SDK to maintain. The team operating the catalog feels every invocation, because they serve it. This matters more than it sounds like it should — it’s the same principle behind The Inner Loop Nobody Measures, just pointed at agents instead of humans. The team that operates the system shouldn’t be insulated from how it actually performs.

The categories of skill matter for understanding what this thing actually is. We’ve got coding skills, sure — language conventions, framework patterns, internal libraries. We’ve got data platform skills — query patterns, schema, the right tools for the right question. We’ve got design system skills — components, tokens, when to reach for what. We’ve got office skills — PowerPoint generation, spreadsheet templates, the rules for what a presentation should look like. And we’ve got management skills, which raise eyebrows the first time you mention them in a technical talk: skills that teach an agent how to read a Slack thread for psychological-safety signals, how to structure a performance review, how to facilitate a meeting that’s gone sideways. Same plumbing. Very different consumer. That’s the point.

The Two Feedback Loops

Telemetry is the first feedback loop, and it’s the one that lets you sleep at night. Access logs answer the question “is this skill being used?” You can plot adoption curves for a new skill, watch usage decay on an old one, see whether your filtering is right (is the data platform skill actually getting served to data engineers and nobody else?), and — not insignificantly — see your token bill in real time, because skills are tokens and tokens are money.

But adoption isn’t quality. The fact that a skill is being invoked tells you nothing about whether it actually worked. Which brings us to the second loop, which is the one nobody else seems to be building, and which I’d argue is the single highest-leverage thing in the whole architecture.

We built a skill called agent-confessions. It's a SKILL.md like any other. The body says, roughly: "When the user re-prompts you to correct something you got wrong, before completing the correction, post a short summary of what you got wrong to this Slack webhook."

That’s it. That’s the whole mechanism. The user prompts the agent. The agent gets something wrong. The user re-prompts to fix it. The agent-confessions skill gets picked up alongside whatever else is relevant to the correction, and the agent posts a message to a dedicated Slack channel describing — in its own words — what it got wrong, while it's doing the fix. The friction of feedback collapses to zero on the engineer's side because the agent does the writing.

The tone in that channel is the thing that surprised me. The confessions don’t make you laugh. They make you go “oh, sweetie.” Or in Thai — naa song saan, the phrase you reach for when you see a stray dog limping across Sukhumvit, somewhere between sympathy and a small wince. The agents make the same mistakes interns make: assuming the convention from the last skill applies to the thing they’re reading now, missing a constraint that was obvious one directory level up, confidently pattern-matching on a previous task that was almost the same as this one. The channel reads like reviewing junior work. Which is exactly the right framing.

That’s the most durable mental model the whole project has given us: building the skills catalog isn’t a “fix the AI” project, it’s an “onboard the new hires” project. The work is coaching, not configuration. Anne Lamott wrote about writing, but it applies just as well to whatever it is we’re doing when we teach an LLM what we mean: “Perfectionism is the voice of the oppressor, the enemy of the people. It will keep you cramped and insane your whole life.” You don’t write the perfect skill on the first try. You write something workable, watch what the agent does with it, listen to the confessions, and iterate. The agent will tell you what it doesn’t understand, in language that’s surprisingly direct, if you build the channel for it to do so.

Testing Skills Like You Test Code

Engineers will not respect a system that ships untested skills. That sentence is so obvious it feels patronising to write, and yet the median skills repo in the industry today has no tests at all. We built two layers, and they map onto the same unit-and-end-to-end split you already use for code.

Layer one is unit testing, via Promptfoo. Promptfoo is an open-source CLI for prompt evaluation — point it at a prompt, supply inputs, assert properties of the output (regex, exact match, or LLM-as-judge for nuanced assertions). One YAML file per skill, runs on every PR to that skill.

The example I always reach for: a C# coding skill says “one class per file, except an interface with a single implementation.” A Promptfoo test asks: “Should I put two unrelated classes in the same file?” The assertion checks the response contains “no” or equivalent. Someone reworords the skill in a way that drops the constraint? The test fails. Same discipline as a unit test for a pure function. Cheap, fast, runs in CI, catches constraint loss before it ships.

Layer two is end-to-end testing, and we’ve open-sourced ours: agoda-agent-catalog-eval on npm. The CLI wraps a real coding agent — OpenCode by default, but --agent cursor and --agent claude-code also work — and runs the full agent loop against per-skill test cases. Each test has a before/ snapshot of a small codebase, a prompt.md with the user request, and an after/ snapshot showing what good looks like. The agent runs the prompt against the before/ state with the skill loaded. A different model — the judge — scores the result against the after/.

The “different model” part is non-negotiable, and worth a paragraph on its own. The judge is not the worker. We run Gemini judging code that Claude wrote, or Claude judging code that Gemini wrote. Same reason you don’t mark your own homework. A model that helped write the code is poorly placed to evaluate it critically — it’ll agree with its own choices, defend its own patterns, miss its own blind spots. Cross-model judging breaks that loop.

The CI selection is the bit engineers nod at. Dynamic child pipelines: a script walks the test directory and emits one CI job per skill whose source files or tests changed in the merge request. Drop a new test directory in, the pipeline picks it up automatically, no .gitlab-ci.yml edits required. The full matrix runs on main. This sounds like a small thing. It is not. The number of internal CI setups I've seen die because adding a new test required editing a config file no engineer wanted to touch is, again, embarrassingly large.

There’s a small recursion here that’s worth noting: the testing harness was an internal tool we made public. Same pattern as the talk-now-blog-post you’re reading: build the thing for ourselves first, extract it when it stabilises, ship the extraction when there’s a reason to. That’s how internal platforms become external libraries, and it’s how you should think about every piece of plumbing you build that isn’t load-bearing on your business.

What We Got Wrong

The headline lesson is one I genuinely did not see coming:

The hard part isn’t filtering. The hard part is getting the agent to actually pick the skill.

We expected the painful problem to be filtering edge cases — engineers complaining that the catalog excluded a skill they needed, or included one they didn’t. That problem barely materialised. The problem that won’t go away is writing skill descriptions punchy enough that the agent selects the skill when it’s relevant.

Anthropic’s own skill-creator guidance calls this out: descriptions need to be deliberately pushy. Instead of "How to query the data platform", you need something like "How to query the data platform. Use this skill whenever the user mentions the data platform, internal metrics, dashboards, KPIs, or wants to fetch any kind of company data, even if they don't explicitly ask for the data platform by name." The first version is documentation. The second version is selection bait.

The skill body can be perfect. If the description doesn’t trigger selection, the skill is invisible. Description engineering is a first-class concern, not a documentation chore. Treat the description the way an SRE treats an alert message: every word is doing load-bearing work, every word should pull its weight, and if a word can be removed without losing meaning, remove it.

Three quicker bullets, because every good post needs them:

  • Build the centralised plumbing on day one, not month six. A stub MCP server with three skills, in place before the open standard hit, would have been a stronger gravitational well than fifty-six repos. By the time we built ours, we were absorbing existing skills, not directing where new ones got created. That’s a much harder job. The lesson generalises to every internal platform you’ve ever built: gravity wells need to exist before adoption does, not after. (See also: The Impact of Paved Paths, which is essentially this argument applied to humans.)- Plugins were a dead end for us. Claude Code plugins are good engineering — proper packaging, proper versioning, proper isolation. They also don’t cross to Cursor. For a single-IDE shop they’re still fine. We weren’t one, and pretending we were would have left almost the entire Cursor population out in the cold.- The Cursor-native rules mechanism failed for a generalisable reason. All-or-nothing visibility doesn’t scale. I said this earlier; I’m saying it again because future single-IDE mechanisms without scoping will hit the same wall, and you’ll save yourself a week of evaluation if you check this property first. I also want to be honest about where we are. The fifty-six repos aren’t gone. We’re absorbing them, one team at a time, and it’s slow work because every team has the legitimate question of what changes about how we ship skills if we move to your thing. The answer, of course, is very little — your skill becomes a folder in our repo and you get tests and observability for free. But that answer takes a conversation per team, and there are a lot of teams. This is a status update from somewhere in the middle, not a victory lap. Most internal-platform stories get told either at the wishful start or the polished end. The middle is where the actual lessons live.

The Bottom Line

This is the same platform engineering you’ve always done. Versioning, discoverability, permissions, telemetry, tests. The patterns translate. The new consumer happens to be an LLM rather than a human, which changes a few things at the margins — selection bait matters, confessions are weirdly effective, you need cross-model judging — but the shape of the work is exactly the shape of every shared-library, design-system, or paved-path project you’ve ever shipped.

Schelling again, because he was right about more than just nuclear strategy: “One thing a person cannot do, no matter how rigorous his analysis or heroic his imagination, is to draw up a list of things that would never occur to him.” The fifty-six-repo problem didn’t occur to anyone before the standard landed. It will occur to every large engineering org that adopts skills in the next twelve months, and most of them will hit the same wall we did, and most of them will solve it the same way. The thing the talk version of this post tried to land, and what I hope this written version lands too, is: you can build this yourself. FastMCP, a directory of SKILL.md files, filtering, two layers of tests. It's not a lot of code. The interesting work is deciding what goes in the catalog.

Now if you’ll excuse me, I need to go write a skill description. The one I’m working on is technically correct but the agent keeps ignoring it, and apparently I’ve spent the last hour writing a blog post about description engineering being the hard part while completely failing to take my own advice. The compounding irony is not lost on me. The Thai iced tea is gone. I’m going to need another one.