An experiment from Causal Map Ltd
An app that answers one evaluation question at a time, and shows its working
Not a coding pass over a pile of documents. You compose an analysis: twelve tests fanned out from one line, a corpus split so the half that wrote a test never scores it, the same instruction run twice on ten sources to measure how much the answer depends on the wording, several tables with their own bases, and a verdict against a rubric somebody registered before any of it ran.
One workflow, and none of it is a batch job
Drawn the way the app draws a workflow in the margin: steps down the left, the assets they produce on the right. Twelve coding prongs come from one line of YAML. The subsample beside them tests the instrument rather than the theory. The two judgements read a rubric that was named, versioned and dated before the first document was opened, and the overall one takes the weakest link rather than the average.
What a batch tool cannot do
Run twenty analyses from one line
A fan-out runs the same steps once per case, per test, per theory or per distinct value in a column. Twenty cases is one line, not twenty copied chains, and the gather afterwards makes a case that turned up nothing appear as a nought instead of vanishing.
Measure its own reliability
Run one instruction twice over the same ten sources and pair the results: the flip rate is how much the answer depends on nothing at all. Run two wordings and the difference is instruction sensitivity. Neither is a finding about the world, and both belong in the report.
Judge against a written standard
A rubric says what the material has to show to earn each verdict, and it is registered with a version, a date and a name. Judgements are tabulated one at a time before anything is rolled up, and the rule for combining them is declared in advance rather than defaulting to an average.
What actually happens
- 1
You point it at documents
A Causal Map project: transcripts, reports, field notes, open-ended survey responses. Text somebody wrote or said, rather than a spreadsheet of numbers.
- 2
You type a question
“Did the mentoring change how young people describe their chances of work?” In your own words, as you would put it to a colleague.
- 3
Ruby drafts a workflow, you argue with it
Ruby is the assistant. She proposes which documents to read, how to mark passages up, how to count the marks and what standard turns a count into a verdict. You change any of it. Nothing has been read yet.
- 4
You register it and run it
The workflow is stamped with a version, a date and your name. Then it codes the documents passage by passage and reports a verdict, a set of figures, and every quote behind them.
And every figure walks back to the words
The four things on the left are the whole agreement, and they are settled before any document is read. Registering them fixes the order rather than the content: revise as often as the work needs, and every version keeps its date. What comes out on the right can be read in either direction, down from the verdict to the words somebody said, or up from a passage to what it ended up counting towards.
Three moves
Break it down
“Was the programme effective?” is four questions wearing one coat: good for whom, compared with what, against whose standard, and how much of that this material can answer. Splitting it is the work, and most of the value is in the split rather than in the answer.
Build it back up
The pieces reassemble into a workflow: which sources to read, how to mark them up, how to count the marks, how to reach a verdict. Simple or elaborate, as the question needs. A two-step workflow is a perfectly good answer to a two-step question.
Keep it checkable
Every figure walks back. Verdict to figures, figure to rows, row to quote, quote to the span in the document. A quotation the machine cannot find in the source is thrown away rather than reported.
A codable question is one that two people would answer much the same way from the same material, with every answer leading back to a passage a third person can check.
How a piece of work goes
- 1
Argue out the question
With Ruby, the assistant. She reads the material, says what she takes the question to mean, and argues with the parts that will not work. She never decides what counts as good.
- 2
Agree the workflow
Sources, coding instruction, counting rule, the standard that turns a count into a verdict. This is the thing you agree, and it is the product.
- 3
Register it, dated
With a version and a name against it. Change a threshold after seeing the numbers and that shows. Change it before and that shows too.
- 4
Run it and read it back
A step at a time if you like: run the coding, look at what came out, fix the instruction, run it again, and only then go on to the counting.
And then play with it as much as you like
None of this is a cage. Try a wording, look at what it marked, throw it away, try another. Run the same instruction twice to see how much the answer moves. Run two rival codings side by side and compare them. Change the model. The discipline is not that you decide once, it is that every version is dated and none is ever rewritten, so a reader can see which side of the evidence each change fell on. Iterate as much as the work needs; what is forbidden is silent change.
Why bother
The reader of an ordinary evaluation report cannot see four things: which passages were counted and which were passed over; what rule set the threshold at “more than half”; whether that rule was written before or after somebody knew how the numbers fell; and whether a second coder would have marked the same passages.
AI makes that worse, because a model reads a thousand pages overnight and nobody can check its working. Used well it raises the quality of work on narrative material more cheaply than anything before. Used badly it produces junk just as fast, and on the page the two look alike: both fluent, both quoting a few passages that fit. The difference falls on whoever acts on the report.
It also makes it fixable, because a machine that reads can be made to show every step.
We claim no determinism and no full reproducibility. AI coding varies between runs, and machine agreement with human coders is well short of perfect. What survives is auditability: a reader who disagrees can find out exactly where.