Wes Ellis./ a personal notebook
Technology. Stories. Side projects.
A few things worth writing down.
← Back to G8KEPR

G8KEPR

A Scorecard Instead of a Vanity Metric

Part 9 of the thread Building G8KEPR

THE SHORT VERSION4 points
  • G8KEPR's features are grouped, each group gets a score from 0.0 to 10.0, and those roll up into one overall number.
  • Score changes stay visible over time, and the differentiators are highlighted so they can't hide in the average.
  • The plan is to publish it at g8kepr.com/scores once it reaches 8.6, then keep raising the bar, like a changelog.
  • Day to day, work lives in Todoist, sorted from P0 (blocks belief or purchase) down to P4 (polish).

Every startup needs a number to look at. The trouble is that the easy numbers are usually the wrong ones. Signups, stars, logos on a slide: they go up and to the right, and they say very little about whether the product is actually good.

For G8KEPR I wanted a measure of the product itself. So I built a scorecard.

How the scorecard works

The structure is simple on purpose:

  1. Features are grouped. Related capabilities sit together, so the score reflects areas of the product rather than a long flat list of checkboxes.
  2. Each group gets a score from 0.0 to 10.0. One decimal place, so small improvements still show up.
  3. The groups roll up into an overall tally. One number for the whole product.
  4. Changes stay visible over time. A group going from, say, 6.2 to 7.0 is information. So is a score that hasn't moved in a month.
  5. Differentiators are highlighted. The groups that make G8KEPR different, above all the cross-pillar correlation engine, are called out so they don't disappear into an average.

That last point matters more than it looks. An overall score can creep upward because a dozen ordinary features got a bit better, while the one thing that makes the product worth buying sits still. Highlighting the differentiators stops the average from flattering me.

What it deliberately leaves out

The scorecard doesn't factor in clients or design partners. Not at all.

That's a deliberate choice, not an oversight. The scorecard answers one question: how good is the product? Who's using it is a different question with its own answers, and mixing the two muddles both. If the product score could go up because of a meeting that went well, it would stop telling me anything about the code.

It also keeps the scorecard honest in a practical way. It can only move when the product changes. There's no shortcut through it that doesn't involve building something.

Publishing it

The plan is to publish the scorecard at g8kepr.com/scores once the overall score reaches 8.6.

After that, the idea is to keep raising the bar, and treat the public page like a changelog. Scores change, the criteria get stricter, and anyone watching can see what moved and when.

Note

It isn't published yet. 8.6 is the threshold I set for putting it in public, and the page doesn't exist until the product earns it.

Why publish at all? Because the scorecard makes a claim about how good the product is, and a claim nobody else can see is easy to fudge, even for the person making it. Putting it where people can read it (and read how it changes) is a way of holding myself to it. It's also a partial answer to the trust question I got into in the open source post. You can't read the code, but you can see how I'm grading it.

P0 to P4: deciding what to work on

A scorecard tells you where you stand. It doesn't tell you what to do on a Tuesday night. That part lives in Todoist.

There's a G8KEPR parent project with about twenty subprojects under it, organized partly by the four pillars: API security, MCP security, the AI gateway and the verification engine. Every subproject is split into the same priority sections:

Section What goes there
P0 — Blocks belief/purchase Anything that would stop someone from believing the product works, or from buying it
P1 In between
P2 In between
P3 In between
P4 — Polish Nice to have, done when everything above it is handled

The labels I care about most are the two ends. P0 is phrased around belief, not just bugs. A feature can work perfectly and still be a P0 problem if nothing about it convinces someone it works. P4 is where polish waits its turn.

How the two fit together

The scorecard and the Todoist lists do different jobs:

  • The scorecard is the scoreboard. It says how the product is doing, group by group, over time.
  • The P0-P4 sections are the playbook. They say what to do next.

Working solo, on evenings and weekends, hours are the scarcest thing I have. Between a score that only rewards real product progress and a list that puts belief-blocking work first, it's harder to spend an evening on something that feels productive and isn't.

When a batch of work is ready to build, it goes from those lists into a three-agent Claude Code workflow. That's the next post.