ENTRY 001
Somebody else’s specs, quoted and checked
Three hundred thirty-three pages of somebody else’s specification, and the only question that mattered was whether it quoted them honestly.
The test was a real, public document: the specs for a school district HVAC replacement,
Pittsburg Unified School District Bid 24-026, 333 pages, issued for bid. Ask and we will
send you the exact PDF.
It read the full mechanical division — all eight Division 23 technical sections, 49 pages of
specification — entirely on the box. Then we took every passage it had quoted and went back
to the source to see whether those words were actually there.
Twelve of twelve quotes check out against the source. Every paragraph
number it cited was real. That is the load-bearing claim of this whole product, and it held
on first contact with a document we did not write.
It surfaced six candidate conflicts and then argued itself out of all six. Every one is
exactly the false positive a keyword matcher would have dropped on a coordinator’s desk:
dielectric unions rated 250 psig against flexible connectors rated 125 psig — different
components, not a contradiction. Duct insulation at 1-1/2" concealed against 1" exposed —
separated by scope. Duct liner against external wrap — different assemblies entirely.
The sixth is the one we keep thinking about. Two sections both call for 5,000 psi grout, and
one of them writes it psig. It declined to call that a conflict and flagged it as
the spec writer’s typo instead, on the grounds that compressive strength is not a gauge
pressure.
None of the six was a real conflict. Six times it declined to waste
our afternoon, and once it caught something the spec writer missed. Precision is the thing
a checking tool has to earn: a list of conflicts too long to read is just a different way of
missing the conflict.
A place to say “fine.” A list of findings needs a legitimate place
to record a negative — we checked this and it’s fine. Without one, a model that
correctly concludes two clauses do not conflict has nowhere to put that conclusion
except the findings list. So the answer has a place for it.
Check the checker. Every quote is compared against the source text. A PDF text layer
can split 10-feet across a line break, and careless whitespace handling turns that
into a mismatch. Strip the whitespace before you accuse a model of making something up.
ENTRY 002
Ask the whole set and it tells you what it read
Every whole-set answer opens with what was and was not read. A bare “not found” hides whether it looked at all.
A question across a whole set works like this: pick the sheets most likely to carry the
answer, read each one, answer with the sheet attached, and open with what was and was not
read. That opening line is the point of the feature.
On the 333-page set from entry 001, every page indexed, the answer opens with
“Searched all 333 pages.”
Search finds the thing you asked for, not its neighbors. “MCC” against that
set returns zero pages, which is the right answer: the specs have no motor
control center in them. A page that merely mentions “motor” is not a hit.
Values are compared only when both carry a unit, and the unit has to be unambiguous.
“3 in the room” and “5 in the corridor” are counts, not three inches
against five, and they are not reported as a conflict.
ENTRY 003
We planted conflicts in a spec. It never called a near miss real.
The public spec in entry 001 turned up no real conflicts. So we wrote one that has them, with traps beside the conflicts, and fixed the scoring rules before the first run.
Entry 001 tests whether a citation is real. Catching a conflict needs a document with a known
answer, so we wrote one: a 38-page invented Division 23 specification with contradictions
planted in it and three near misses that look like conflicts and are not.
One planted pair is the classic: a cross-reference that repeats an insulation thickness
instead of pointing at it, and the value went stale when the other section was revised.
The scoring rules were written down before the first run. A near miss returned as a real
conflict is a false alarm.
Five runs. No near miss was called real in any run, and neither was any
of the eleven pairs the model raised on its own.
So the pair of numbers is this. Citation: twelve of twelve, on somebody else’s real
document. False alarms: none, in any run, on a specimen written to set them off.
ENTRY 004
Graded against a hand-checked sheet, not against itself
A status of ok is not the same as correct. So every change to how it reads schedules is scored against sets checked by hand, and a gate decides whether it stays.
A tool that grades itself on whether it found a schedule and how many rows it got out of it
can only go up. A tool that reads a general note as a table looks, by that measure, like a
tool that got better. Row counts are not how we score.
So we hand-check. Take a real sheet, work out what the table actually says, check every mark
and every dimension against the printed page, and score the tool against that instead of
against itself. One school district set has a door schedule and a window schedule, one page
each, hand-checked mark by mark.
Every change to how it reads a schedule is scored against six hand-checked
sets. A change that makes documents worse is refused by the gate, however good its
headline number, and what stands is the last build that passed with nothing regressed.
The whole argument for this product is that its answers are checkable. A score we can
overrule is not a check.
ENTRY 005
Every public set we could find, each counted once
Sets drawn by people who had never heard of us, and one command that scores every one of them against a fixed baseline.
We went and got real sets: public drawing sets pulled from school district bid portals, city
permit portals, university facilities archives, and the free preapproved accessory-dwelling
plans that several California counties publish. Drawn by real people who had never heard of
us.
Every file in the collection is identified by what is in it, not by what somebody named it.
Counted that way: 218 files, 184 genuinely distinct documents.
Thirty-four were exact copies — one file had been downloaded four separate times.
Seven were web error pages saved with a .pdf name. Everything is scored
against real drawing sets, each counted once.
Counting a document by its contents rather than its filename sounds like housekeeping. It is
the difference between “this got better” and “this got counted twice.”
Every drafting office formats a schedule differently. Headers that stack across two lines,
width and height in separate columns, schedules with no orientation column at all, and one
office’s habit of setting the inches of 2'‑0" on a slightly different
baseline from the feet. Teaching the reader one of those is easy to do on the set in front of
us. The question is what it does to all the others.
So the scoring is the point of this entry, not any one format. One command scores every set
against a fixed baseline, and a change that helps the document in front of us but costs
others shows up as a number instead of a good feeling. New schedule formats go in one at a
time, each proved against the whole collection before it stays.
ENTRY 006
The model reads. Code does the arithmetic.
The first end-to-end run set the rule the calculation engine is built on.
We ran the thing end to end for the first time: hand it a PDF of building inputs, have it
pull out the details, and calculate from them. Every figure it produced was arithmetically
correct.
The rule that came out of it: the model reads the
document, and vetted code does the math. A language model is good at finding
“R-21” buried in a general note. Deciding what goes into a calculation, and doing
it, is a job for code that can be checked. Those are different jobs, and they stay separate.
ENTRY 007
One box, and every ceiling was software
One machine, 128 GB of memory that its processor and graphics chip share, and every limit we hit turned out to be software. No parts were ordered, and nothing was replaced.
The box is a single AMD Ryzen AI Max+ 395. Its 128 GB is one pool, shared between the processor
and the graphics chip, and the split between them is a setting. Set that split badly and you
have bought memory you cannot use.
There were three walls, and none of them were made of silicon: the memory split, the program
that runs the model, and the way requests were packaged. No part was changed to get past any
of them.
The practical result is that one model now handles text and images together.