An experiment from Causal Map Ltd

An app that answers one evaluation question at a time, and shows its working

Not a coding pass over a pile of documents. You compose an analysis: twelve tests fanned out from one line, a corpus split so the half that wrote a test never scores it, the same instruction run twice on ten sources to measure how much the answer depends on the wording, several tables with their own bases, and a verdict against a rubric somebody registered before any of it ran.

How it works

One workflow, and none of it is a batch job

A workflow that is not a straight line Sampling all fifty sources with a seed, splitting them into a design half and a held-out half, fanning twelve coding prongs out over the held-out half at a set chunk size, aligning a repeated coding on a ten-source subsample to measure stability, counting shares per test with the base attached, and judging twice against a registered rubric, per test and then overall with the weakest link deciding. DRAW THE FRAME SAMPLE All fifty sources method: all · seed 4291 frame 50 units, the base every later figure states SPLIT IT, AND KEEP THEM APART SELECT Design half, held-out half the half that wrote a test never scores it two halves, 25 each the run records which sources it read ONE PRONG PER TEST CODE hypervigilance 1 CODE displacement 2 CODE … test 12 for_each: {values_of: test} · chunk 2,000 words · segment quota on · every prong the same budget CHECK THE INSTRUMENT ALIGN The same instruction twice on a ten-source subsample · run-to-run stability pairs one row per pair, an agree_ flag per column COUNT, WITH A STATED BASE COMPUTE Share of units, per test over: {group: item} · and again where no speaker delivers two tables twelve rows each, denominator on every row JUDGE TWICE, ON TWO RUBRICS JUDGE Per test, then overall rubric loneliness-tests v3 · combine: worst 12 judgements verdict Every box is a few lines of YAML you can argue with before it runs, and every arrow is an asset a reader can open.

Drawn the way the app draws a workflow in the margin: steps down the left, the assets they produce on the right. Twelve coding prongs come from one line of YAML. The subsample beside them tests the instrument rather than the theory. The two judgements read a rubric that was named, versioned and dated before the first document was opened, and the overall one takes the weakest link rather than the average.

What a batch tool cannot do

Run twenty analyses from one line

A fan-out runs the same steps once per case, per test, per theory or per distinct value in a column. Twenty cases is one line, not twenty copied chains, and the gather afterwards makes a case that turned up nothing appear as a nought instead of vanishing.

Measure its own reliability

Run one instruction twice over the same ten sources and pair the results: the flip rate is how much the answer depends on nothing at all. Run two wordings and the difference is instruction sensitivity. Neither is a finding about the world, and both belong in the report.

Judge against a written standard

A rubric says what the material has to show to earn each verdict, and it is registered with a version, a date and a name. Judgements are tabulated one at a time before anything is rolled up, and the rule for combining them is declared in advance rather than defaulting to an average.

What actually happens

  1. 1

    You point it at documents

    A Causal Map project: transcripts, reports, field notes, open-ended survey responses. Text somebody wrote or said, rather than a spreadsheet of numbers.

  2. 2

    You type a question

    “Did the mentoring change how young people describe their chances of work?” In your own words, as you would put it to a colleague.

  3. 3

    Ruby drafts a workflow, you argue with it

    Ruby is the assistant. She proposes which documents to read, how to mark passages up, how to count the marks and what standard turns a count into a verdict. You change any of it. Nothing has been read yet.

  4. 4

    You register it and run it

    The workflow is stamped with a version, a date and your name. Then it codes the documents passage by passage and reports a verdict, a set of figures, and every quote behind them.

And every figure walks back to the words

How a workflow is put together and read back Four agreed steps, which sources to read, how to mark passages, how to count the marks and what counts as good; these are registered with a date and a name; running them produces a verdict whose figures, coded rows and quotes lead back to the words in the document. WHAT YOU AGREE, BEFORE ANYTHING IS READ 1 Which sources to read All of them, or a stated sample 2 How to mark passages up The coding instruction, in writing 3 How to count the marks Per person, per passage, per source 4 What would count as good The rubric, and who put their name to it REGISTER dated, versioned, with a name on it then run it WHAT COMES OUT, AND WHAT IT RESTS ON A verdict Figures Coded rows Quotes The words somebody said walks back to A place where a person looks, changes it, and runs it again

The four things on the left are the whole agreement, and they are settled before any document is read. Registering them fixes the order rather than the content: revise as often as the work needs, and every version keeps its date. What comes out on the right can be read in either direction, down from the verdict to the words somebody said, or up from a passage to what it ended up counting towards.

Three moves

Break it down

“Was the programme effective?” is four questions wearing one coat: good for whom, compared with what, against whose standard, and how much of that this material can answer. Splitting it is the work, and most of the value is in the split rather than in the answer.

Build it back up

The pieces reassemble into a workflow: which sources to read, how to mark them up, how to count the marks, how to reach a verdict. Simple or elaborate, as the question needs. A two-step workflow is a perfectly good answer to a two-step question.

Keep it checkable

Every figure walks back. Verdict to figures, figure to rows, row to quote, quote to the span in the document. A quotation the machine cannot find in the source is thrown away rather than reported.

A codable question is one that two people would answer much the same way from the same material, with every answer leading back to a passage a third person can check.

From the principles. Most questions arrive uncodable, and making them codable is the job.

How a piece of work goes

  1. 1

    Argue out the question

    With Ruby, the assistant. She reads the material, says what she takes the question to mean, and argues with the parts that will not work. She never decides what counts as good.

  2. 2

    Agree the workflow

    Sources, coding instruction, counting rule, the standard that turns a count into a verdict. This is the thing you agree, and it is the product.

  3. 3

    Register it, dated

    With a version and a name against it. Change a threshold after seeing the numbers and that shows. Change it before and that shows too.

  4. 4

    Run it and read it back

    A step at a time if you like: run the coding, look at what came out, fix the instruction, run it again, and only then go on to the counting.

And then play with it as much as you like

None of this is a cage. Try a wording, look at what it marked, throw it away, try another. Run the same instruction twice to see how much the answer moves. Run two rival codings side by side and compare them. Change the model. The discipline is not that you decide once, it is that every version is dated and none is ever rewritten, so a reader can see which side of the evidence each change fell on. Iterate as much as the work needs; what is forbidden is silent change.

Why bother

The reader of an ordinary evaluation report cannot see four things: which passages were counted and which were passed over; what rule set the threshold at “more than half”; whether that rule was written before or after somebody knew how the numbers fell; and whether a second coder would have marked the same passages.

AI makes that worse, because a model reads a thousand pages overnight and nobody can check its working. Used well it raises the quality of work on narrative material more cheaply than anything before. Used badly it produces junk just as fast, and on the page the two look alike: both fluent, both quoting a few passages that fit. The difference falls on whoever acts on the report.

It also makes it fixable, because a machine that reads can be made to show every step.

We claim no determinism and no full reproducibility. AI coding varies between runs, and machine agreement with human coders is well short of perfect. What survives is auditability: a reader who disagrees can find out exactly where.

It is early, and published early on purpose

Nothing here has been run on a live evaluation. The arguments are drafted as working papers so people can disagree with them before there are results to defend.

Read the working papers