Build log

What we built and how we check it, one entry at a time.

Grouped in three parts: what it reads, how we check it, and how it is built. Every figure here was measured on one machine — an AMD Ryzen AI Max+ 395 with 128 GB of memory that its graphics chip shares — and every run described is local.

ENTRY 001

Somebody else’s specs, quoted and checked

Three hundred thirty-three pages of somebody else’s specification, and the only question that mattered was whether it quoted them honestly.

The test was a real, public document: the specs for a school district HVAC replacement, Pittsburg Unified School District Bid 24-026, 333 pages, issued for bid. Ask and we will send you the exact PDF.

It read the full mechanical division — all eight Division 23 technical sections, 49 pages of specification — entirely on the box. Then we took every passage it had quoted and went back to the source to see whether those words were actually there.

Twelve of twelve quotes check out against the source. Every paragraph number it cited was real. That is the load-bearing claim of this whole product, and it held on first contact with a document we did not write.

It surfaced six candidate conflicts and then argued itself out of all six. Every one is exactly the false positive a keyword matcher would have dropped on a coordinator’s desk: dielectric unions rated 250 psig against flexible connectors rated 125 psig — different components, not a contradiction. Duct insulation at 1-1/2" concealed against 1" exposed — separated by scope. Duct liner against external wrap — different assemblies entirely.

The sixth is the one we keep thinking about. Two sections both call for 5,000 psi grout, and one of them writes it psig. It declined to call that a conflict and flagged it as the spec writer’s typo instead, on the grounds that compressive strength is not a gauge pressure.

None of the six was a real conflict. Six times it declined to waste our afternoon, and once it caught something the spec writer missed. Precision is the thing a checking tool has to earn: a list of conflicts too long to read is just a different way of missing the conflict.

A place to say “fine.” A list of findings needs a legitimate place to record a negative — we checked this and it’s fine. Without one, a model that correctly concludes two clauses do not conflict has nowhere to put that conclusion except the findings list. So the answer has a place for it.

Check the checker. Every quote is compared against the source text. A PDF text layer can split 10-feet across a line break, and careless whitespace handling turns that into a mismatch. Strip the whitespace before you accuse a model of making something up.

ENTRY 002

Ask the whole set and it tells you what it read

Every whole-set answer opens with what was and was not read. A bare “not found” hides whether it looked at all.

A question across a whole set works like this: pick the sheets most likely to carry the answer, read each one, answer with the sheet attached, and open with what was and was not read. That opening line is the point of the feature.

On the 333-page set from entry 001, every page indexed, the answer opens with “Searched all 333 pages.”

Search finds the thing you asked for, not its neighbors. “MCC” against that set returns zero pages, which is the right answer: the specs have no motor control center in them. A page that merely mentions “motor” is not a hit.

Values are compared only when both carry a unit, and the unit has to be unambiguous. “3 in the room” and “5 in the corridor” are counts, not three inches against five, and they are not reported as a conflict.

ENTRY 003

We planted conflicts in a spec. It never called a near miss real.

The public spec in entry 001 turned up no real conflicts. So we wrote one that has them, with traps beside the conflicts, and fixed the scoring rules before the first run.

Entry 001 tests whether a citation is real. Catching a conflict needs a document with a known answer, so we wrote one: a 38-page invented Division 23 specification with contradictions planted in it and three near misses that look like conflicts and are not. One planted pair is the classic: a cross-reference that repeats an insulation thickness instead of pointing at it, and the value went stale when the other section was revised.

The scoring rules were written down before the first run. A near miss returned as a real conflict is a false alarm.

Five runs. No near miss was called real in any run, and neither was any of the eleven pairs the model raised on its own.

So the pair of numbers is this. Citation: twelve of twelve, on somebody else’s real document. False alarms: none, in any run, on a specimen written to set them off.

ENTRY 004

Graded against a hand-checked sheet, not against itself

A status of ok is not the same as correct. So every change to how it reads schedules is scored against sets checked by hand, and a gate decides whether it stays.

A tool that grades itself on whether it found a schedule and how many rows it got out of it can only go up. A tool that reads a general note as a table looks, by that measure, like a tool that got better. Row counts are not how we score.

So we hand-check. Take a real sheet, work out what the table actually says, check every mark and every dimension against the printed page, and score the tool against that instead of against itself. One school district set has a door schedule and a window schedule, one page each, hand-checked mark by mark.

Every change to how it reads a schedule is scored against six hand-checked sets. A change that makes documents worse is refused by the gate, however good its headline number, and what stands is the last build that passed with nothing regressed.

The whole argument for this product is that its answers are checkable. A score we can overrule is not a check.

ENTRY 005

Every public set we could find, each counted once

Sets drawn by people who had never heard of us, and one command that scores every one of them against a fixed baseline.

We went and got real sets: public drawing sets pulled from school district bid portals, city permit portals, university facilities archives, and the free preapproved accessory-dwelling plans that several California counties publish. Drawn by real people who had never heard of us.

Every file in the collection is identified by what is in it, not by what somebody named it. Counted that way: 218 files, 184 genuinely distinct documents. Thirty-four were exact copies — one file had been downloaded four separate times. Seven were web error pages saved with a .pdf name. Everything is scored against real drawing sets, each counted once.

Counting a document by its contents rather than its filename sounds like housekeeping. It is the difference between “this got better” and “this got counted twice.”

Every drafting office formats a schedule differently. Headers that stack across two lines, width and height in separate columns, schedules with no orientation column at all, and one office’s habit of setting the inches of 2'‑0" on a slightly different baseline from the feet. Teaching the reader one of those is easy to do on the set in front of us. The question is what it does to all the others.

So the scoring is the point of this entry, not any one format. One command scores every set against a fixed baseline, and a change that helps the document in front of us but costs others shows up as a number instead of a good feeling. New schedule formats go in one at a time, each proved against the whole collection before it stays.

ENTRY 006

The model reads. Code does the arithmetic.

The first end-to-end run set the rule the calculation engine is built on.

We ran the thing end to end for the first time: hand it a PDF of building inputs, have it pull out the details, and calculate from them. Every figure it produced was arithmetically correct.

The rule that came out of it: the model reads the document, and vetted code does the math. A language model is good at finding “R-21” buried in a general note. Deciding what goes into a calculation, and doing it, is a job for code that can be checked. Those are different jobs, and they stay separate.

ENTRY 007

One box, and every ceiling was software

One machine, 128 GB of memory that its processor and graphics chip share, and every limit we hit turned out to be software. No parts were ordered, and nothing was replaced.

The box is a single AMD Ryzen AI Max+ 395. Its 128 GB is one pool, shared between the processor and the graphics chip, and the split between them is a setting. Set that split badly and you have bought memory you cannot use.

MeasuredBeforeAfter
Usable graphics memory58.8 GB120 GB
Hardware changed—None

There were three walls, and none of them were made of silicon: the memory split, the program that runs the model, and the way requests were packaged. No part was changed to get past any of them.

The practical result is that one model now handles text and images together.

Send a set you already know 

Redgorge reads a full construction set on a box in your building and cites the sheet for every answer. This page is the record of how it was built and how it is tested. We sell the work around it, one fixed-price task at a time — prices are on the front page. The fairest test is a closed-out job you already know the answers to: free, results inside five business days.