An AI assessment for small business is supposed to tell you where automation will pay off and where it won't. Most of them don't do that. They hand back a score built on something you can't check, dressed up as insight. The Blueprint works the other way: we score what has evidence behind it, and we leave the rest blank until it earns one. Here are the seven areas we score, why the blanks matter as much as the numbers, and how to read anyone else's scorecard once you know what to ask.
The seven areas an AI assessment for small business actually scores
Systems and data access. Can the AI actually reach the software you run on, and under what permissions. The evidence is a live check of each system's API or connector, whether access is read-only or read-and-act, and whether the account credentials that would be used actually work. Not a guess about "most platforms." A check of yours.
Task volume and repetition. How much of the week goes to work that repeats the same shape every time. The evidence is either a time log the owner already keeps, a walkthrough of the task with the person who does it, or a count pulled from the software itself — number of quotes sent, calls logged, forms filed. If none of those exist, this area doesn't get scored high; it gets flagged as needing a log first.
Data quality. Automation built on bad data produces bad output faster. The evidence is a sample pull from the actual records — duplicate customers, missing fields, inconsistent formats — checked by hand, not assumed clean because the software looks tidy.
Existing automation. What's already been tried, and whether it's still running or quietly broken. The evidence is an inventory of current Zapier rules, macros, scheduled reports, and scripts, tested to see which ones still fire.
Process documentation. Is there a written process to automate, or does the whole thing live in one person's head. The evidence is whether a step-by-step exists on paper, or whether the person doing the work can walk through it start to finish without skipping steps they didn't realize they were doing.
Permission and compliance constraints. Some data can't leave a system, or can't be touched without review, for reasons that have nothing to do with technology. The evidence is a look at the platform's terms of service and any industry-specific restriction that applies, not a blanket assumption that it's fine.
Cost of the status quo. What the manual version is costing right now, in hours and dollars, so the fix can be measured against something real. The evidence is the arithmetic, worked with the owner: minutes per task, times how often it happens, times a loaded hourly rate. Say a shop logs forty invoices a week at six minutes each, and the person doing it costs the business $28 an hour loaded. That's roughly four hours a week, or about $112 — call it $5,800 a year, before anything else changes. That number comes from the shop's own count, not an industry average.
Each of those seven gets a score from one to five. Every score has a line behind it that says what it's based on: a screenshot, a log, a tested connection, a documented process. If we can't produce that line, the area doesn't get a number.
Why some boxes stay blank
A Blueprint report will sometimes come back with two or three areas marked unscored instead of rated. That's not a gap in the work. It's the honest outcome when the evidence doesn't exist yet.
Say a business wants a score on "employee adoption risk" — how likely the team is to actually use whatever gets built. There's no way to measure that before the system exists and people have tried it. Scoring it anyway means inventing a number to fill a row, and that number will get treated like the rest of the report: as something solid enough to make a decision on.
Same with vendor roadmap reliability. If a platform's API is scheduled to change in six months, nobody outside that company knows for certain, including us. A guess dressed up as a score just moves the risk downstream, onto the reader who trusted the report.
The blank stays visible on the page. It doesn't get folded into an average or hidden behind a composite score. That's the whole point of separating the seven scored areas from everything else: a reader should be able to see exactly where the report ran out of evidence, instead of a single tidy number that quietly absorbed the gap.
This is also why we don't publish an overall "readiness score" as one figure. Averaging five real scores with two guesses produces a number that looks precise and means nothing. Seven separate lines, each with its own backing, tell you more than one blended score ever could.
How to read anyone's scorecard
You don't need our framework to check somebody else's business AI readiness scoring. You need one question, asked seven times: what sits behind this number.
If the answer is a log, a test, a document, or a count from the actual software, that's a score you can trust as far as the evidence goes. If the answer is "our methodology" or "based on our experience with similar businesses," that's not evidence. That's a name for a guess.
Ask the same question about anything left blank, if there is anything left blank. If every single row on a scorecard is filled in, on the first pass, before anyone has looked at your systems, that's worth noticing too. Full marks everywhere usually means nobody checked, not that everything checked out.
An honest AI audit will have some rows it can't fill yet. It will say so, and it will tell you what would need to happen — a log kept for two weeks, an API tested, a process written down — before that row gets a real answer. That's not a weaker report. It's the only kind you can act on.
The AI scorecard itself is just a page. What makes it useful is whether every number on it can be traced back to something you could go check yourself.

