Recursive Systems

Metaphorum 2026, Alliance Manchester Business School

From Evidence to Inquiry: Keeping LLM Inference Provisional in Recursive VSM Diagnosis

When should AI-proposed structure become a diagnostic premise?

Open the frozen case →

Read-only, as the application showed this case on 17 September 2026.

Start with three situations one label would make identical.

Situation What the diagnosis must preserve
A committee is required to review performance. Assigned responsibility.
Minutes show it performed a review. Reported enactment.
Evidence shows the review changed a decision. A working regulatory relationship, with an observable consequence.

Calling all three the same thing discards consequential distinctions.

The route at a glance

Documents and answers Extract entities, relations, quotes Assemble candidates Audit: classify or abstain Map, level by level Diagnosis: findings and questions A person decides rules model diagnosis first function-led inquiry

Every box is a stage on this page; the colours match the route map below.

Recorded in the case. Included in this model request. Admissible for this particular claim. Accepted into the working representation. Selected for the final report. Those are five different things and the system keeps them apart.

This page is a declaration of how the analysis works: I am one person with a prototype, and the parts that are not built yet are named here rather than left for you to discover.

The full declaration, every stage's inputs and outputs and the gate's 28 grounds, is what the product tests itself against, and is shown to pilot partners.

I am not promising a diagnosis, and I am not promising that the first run will be useful. What I can promise is that you will see where every claim came from, and that your corrections will be recorded as yours.

If you have a view on which kind of case would break this method fastest, I would like to hear it. That is more useful to me right now than a case it would handle well.

Read the abstract, as submitted on 28 July 2026

Purpose. Clemson’s VSM ToolBox and Pérez Ríos’s VSMod showed software can support Stafford Beer’s Viable System Model; however, LLMs introduce a cybernetic problem: plausible inference can harden into organisational fact and propagate across recursion levels. This project traces a path from evidence to hypotheses, questions and interventions.

Methods. The prototype ran on two public-record cases: Robin Hood Energy, and the Charity Commission inquiry into New Wineskins and U-Turn Move on Homes. It proposes operations, relations and recursive structure, promoting each only when an evidence-gated test passes, abstaining otherwise. An operation becomes a child system one level down only when at least two cited source units evidence its internal operations and boundary. Structural gaps become focused questions; answers and corrections persist through reanalysis. Hypotheses and prescriptions carry checks that could confirm or refute them. Any LLM-attributed quotation must match its cited source verbatim.

Findings. Both produced evidence-linked maps and explicit gaps. In one, an entered answer that no role monitored the external environment or future was retained verbatim as evidence of a System 4 gap; no function was invented. VSM operates as a reasoning grammar, not only a diagram.

Implications/Contributions. The contribution is an evidence-to-inquiry loop that turns abstention into diagnostic work while model suggestions remain provisional, complementing Metaphorum 2025’s work on AI-augmented VSM decision environments. Results demonstrate software traceability, not diagnostic validity or organisational impact; practitioner testing comes next. The session is a live demonstration: participants answer gap questions and see reanalysis. When should AI-proposed structure become a diagnostic premise?

Update, 6 August 2026: Robin Hood Energy was withdrawn after cold probes showed the models already knew the case's history.

What this is, and is not

Model output is provisional unless a recorded decision says otherwise. That rule is the whole design. A language model reading a transcript will happily name a department a viable system, and once it has, everything downstream inherits the mistake.

Beer names that exact error. "the most common mistake is to seize on the existing organization chart of the institution, and blithely to assume that every division or department shown as depending from the boss is a viable system in its own right" (The Heart of Enterprise). That sentence is the reason the abstentions exist. The machine is faster at that mistake than any consultant, so it has to be stopped at the point of typing, not corrected afterwards.

Three commitments follow from it.

Every promoted claim quotes the passage it rests on. The quotation is checked back against the source, character for character, before the claim is allowed to stand.

Where the evidence does not establish the function, the analysis abstains and records why. An abstention is a piece of diagnostic work, not a failure. It becomes a question for a person.

A quotation that matches its source shows where an interpretation came from. It does not show that the interpretation is right. Beer put the limit plainly: "a model is neither true nor false: it is more or less useful" (Diagnosing the System for Organizations, 1985, pages 1 to 2).

What this is not. It is not a diagnosis, and nothing here has been validated with the people in any organisation described. It is not a measure of how well an organisation works. It is not a finished product.

Where the analysis commits itself

How the analysis routes differ

The four routes side by side

Rules assemble the candidates Identity is settled by exact label bytes and entity kind, before anything is read for function.
Its distinctive commitment
Identity is settled by exact label bytes and entity kind, before anything is read for function.
What it enables
A repeatable candidate set, and a closed per-candidate evidence allowance a person can audit.
What it can attenuate or lose
Two spellings of one team stay apart, unknown kinds are not merged, and a relevant passage can sit unlinked.
Recursive consequence
A split identity cannot become a system in focus one level down, because only a surviving candidate is considered.
A model assembles the candidates Identity is a model interpretation, made before any functional question.
Its distinctive commitment
Identity is a model interpretation, made before any functional question.
What it enables
Mentions that no rule would join can be proposed as one body.
What it can attenuate or lose
The framing is reduced to identifiers, level and parent, and the boundary statement is not in it.
Recursive consequence
The identity that may later be expanded was itself proposed by a model, and the recursion inherits that.
Diagnosis first An account of the case comes first, and structure is projected from it.
Its distinctive commitment
An account of the case comes first, and structure is projected from it.
What it enables
Alternative mechanisms, risks and disconfirming checks are invited, and the account is kept alongside the structure.
What it can attenuate or lose
Prior structure is shown in a capped view, and an account can be grounded at every quotation and still be wrong.
Recursive consequence
Only qualifying candidates are projected, so a hypothesis can be retained while contributing nothing below.
Function-led inquiry The question is fixed and the evidence basis is sealed for the run.
Its distinctive commitment
The question is fixed and the evidence basis is sealed for the run.
What it enables
The model chooses what to look at next, and a contested interpretation stays reviewable.
What it can attenuate or lose
Budgets bound the search, working memory resets between functions, and findings must bind to captured identities.
Recursive consequence
Nothing enters the map until a person promotes it, which queues an ordinary map run.
Route Its distinctive commitment What it enables What it can attenuate or lose Recursive consequence
Rules assemble the candidates Identity is settled by exact label bytes and entity kind, before anything is read for function. A repeatable candidate set, and a closed per-candidate evidence allowance a person can audit. Two spellings of one team stay apart, unknown kinds are not merged, and a relevant passage can sit unlinked. A split identity cannot become a system in focus one level down, because only a surviving candidate is considered.
A model assembles the candidates Identity is a model interpretation, made before any functional question. Mentions that no rule would join can be proposed as one body. The framing is reduced to identifiers, level and parent, and the boundary statement is not in it. The identity that may later be expanded was itself proposed by a model, and the recursion inherits that.
Diagnosis first An account of the case comes first, and structure is projected from it. Alternative mechanisms, risks and disconfirming checks are invited, and the account is kept alongside the structure. Prior structure is shown in a capped view, and an account can be grounded at every quotation and still be wrong. Only qualifying candidates are projected, so a hypothesis can be retained while contributing nothing below.
Function-led inquiry The question is fixed and the evidence basis is sealed for the run. The model chooses what to look at next, and a contested interpretation stays reviewable. Budgets bound the search, working memory resets between functions, and findings must bind to captured identities. Nothing enters the map until a person promotes it, which queues an ordinary map run.

These are design consequences and risks to test, not a measured ranking.

I am comparing where the analysis makes its commitments.

Two requisite-variety questions sit behind that. Does the organisation have the variety its environment demands, and does the inquiry have the variety the organisation demands? The software supports the inquiry's variety in order to investigate the organisation's. It has not demonstrated that either has satisfied requisite variety.

What the observer can distinguish For Ashby, variety counts distinguishable elements, and the count depends on the observer's powers of discrimination.

Each route is an observer with limited powers, fixed before any label is applied.

The corpus is attenuated before it is read. Spans are sorted by a content-stable hash and cut into windows, twelve at a time by default. A window is read with the structure accumulated from earlier windows, not with the whole consulting conversation. The coverage check confirms the spans were put in front of the model, not that any relationship was understood. That is not a fault: the corpus carries more variety than any reader, and what matters is which distinctions survive.

Rules assemble the candidates The first route discriminates by exact label bytes and known entity kind.

It is a crude discriminator, chosen because it repeats and can be checked. No model call is made there, though the reading before and the assessment after are model work. The assessment sees only what one candidate's own evidence allows: its grounding followed back into the extraction records, plus one hop along recorded relationships, with the direction of each relation kept, so being named as the object of an activity is not evidence of performing it. The model may see a passage selected for another candidate in the batch and may not cite it here; output that escapes the allowance is dropped.

Inside that allowance the prompt asks three separate things. Does the source show an assigned responsibility, a reported practice, or an observed performance. A label settles none of them. Where evidence fits several functions the gate refuses to pick one, because that would destroy the distinction. For channels the software records amplification, attenuation and transduction in each direction, and requires no measured capacity.

A decomposition holds while it preserves the connectivity inside each chunk, which is Espejo's warning about fragmentation. A byte test does not know what it is cutting.

A model assembles the candidates instead The second route hands identity to the model.

From the source units, the extraction pointers and the prior structure, it proposes which mentions belong together. It applies no functional labels, and new candidates are untyped. Identity is therefore an interpretation made before any functional question, and what follows inherits it.

The diagnosis comes first The third route holds back less.

It hands over the corpus, the declared frame and a capped view of the prior structure. The prompt asks for alternative mechanisms, risks and disconfirming checks, which widens the account rather than narrowing it. Afterwards the code separates a source observation from the model's interpretation, and checks every quotation backwards against the sources. A status the model declares for itself authorises nothing.

The function asks for its own evidence The fourth route fixes the question instead of the candidate, across six functional inquiries over a basis sealed for the run.

The model searches, fetches passages and revises its account over successive turns, choosing what to look at next. Budgets bound the search, and working memory starts fresh at each function, so what the model learned in one does not carry into the next.

Recursion changes the system in focus

Going down a level is not zooming in. It changes what is being observed. A child is considered only for an operation with a supported contribution, after a viability question and a grounded criterion. At least two source units are required. The criterion is not the level. Pérez-Ríos separates the criterion you unfold along from the level you reach by unfolding. Choose a different criterion and the same organisation unfolds into a different set of levels. The software holds one criterion at a time, declared by a person.

What a person contributes

A gap the analysis cannot close becomes a question, which is Pangaro's loop. The next question sets the goal for the next conversation, and the people in the last one are only possible participants. The software produces the question; who should answer it is a judgement it does not make.

The system keeps separate records for an answer, an identity correction, an analytical review, and a promotion. A review carries its own provenance and can be superseded. Promotion is a separate human command that queues an ordinary map run. A disagreement can be recorded in a person's own words, with no model reply and no extraction over them.

Two provenance views Forward: the review journey follows one recorded review into every later run, showing whether the retained request held the review's text or only its id.
  • Forward: the review journey follows one recorded review into every later run, showing whether the retained request held the review's text or only its id. The application shows that list.
  • Backward: the output trace takes one output to its stages and retained requests, and answers "unavailable" rather than guessing. The application shows that route.
  • Neither view claims influence. A shared source unit is not a shared passage. The application states that limit on its own reading.
The full route map

How the analysis paths connect

Read downward. Each shared preparation stage spans all three columns; each shared Map stage spans the two Map columns. Individual stages stay in their own lane. Open any stage to see what it receives, does and produces.

The two Map columns group three fresh-seed choices within the same standard Map strategy: code skeleton, LLM skeleton, or integrated inverted seed. A Map can also be asked for factual preparation alone, which stops after the deterministic skeleton. Function-first has its own capture and execution, and can finish with findings without Standard Map Audit. Its reviewed roles enter a new standard Map through a separate human promotion command. Diagnosis is queued separately again, after a Map has settled.

Where a stage depends on a setting, the setting is described in words and the page says which way it is set in the shipped configuration. Optional stages, explicit resume and inquiry branches are labelled. What executed on a particular case is established by that run's own record.

Grey: shared source preparation Blue: skeleton seed Orange: inverted seed Purple: Function-first Blue + orange: both Map modes Teal: Diagnosis, a separately queued run Green: human action
Map: skeleton seedDeterministic or LLM
Map: inverted seedA deep run, unless the flow is switched off
Function-firstSeparate assessment run
Document ingestion and parsingShared by all three paths
Receives
Uploaded documents
Does
Parse content into stored source spans; extraction runs only when requested.
Produces
Stored document spans
Document extraction: reasoningWindowed extraction; repeats over longer documents
Receives
Document spans in windows
Does
Produce a prose evidence ledger; carry prior structure to the next window.
Produces
Prose ledger
Document extraction: formattingSeparate from extraction reasoning
Receives
Prose evidence ledger
Does
Format and resolve extracted nodes against the document spans.
Produces
Resolved nodes and source references
Relation recoveryOnly when the formatting pass resolved two or more entities and no relation at all
Receives
Resolved nodes and spans
Does
Recover document-local relationships for emission when that condition holds. When it does not, the stage records that it was not needed and no model is called.
Produces
Pending relation emissions, or a recorded “not needed”
Document channel signalsOptional fourth stage of the same attempt, after relation recovery; off unless the operator turns it on
Receives
The window's document spans
Does
When the flag is on, extract grounded channel signals in a separate model call with its own budget. With the flag off the attempt is byte-identical to the producer before the feature existed.
Produces
Signals with verbatim source phrases
Evidence emissionJoins extraction, relations and optional signals
Receives
Resolved extraction, relations and optional signals
Does
Emit document evidence and preserve source links; stored evidence can be selected for later analysis.
Produces
Document evidence events and their source spans
Framing declarationShared case frame; independent of ingestion order
Receives
System in focus, recursion criterion, TASCOI and presenting concern
Does
A person declares the frame used for analysis. The selected frame and sources feed each later run; source bytes remain evidence.
Produces
Recorded framing
Map run scope and the analytical basisBoth Map columns use the standard Map strategy
Receives
Frame, scoped evidence and chat pointers
Does
A queued run carries its scope; no stage selects inputs before skeletonisation, which is the first operation. The values the analysis actually consumed are captured inside the reasoning stage, before its model call, and are bound to the run only when settlement succeeds. A failed run certifies no basis.
Produces
A run-scoped basis on success, nothing on failure
Function-first preparation and own snapshotOff unless the operator turns it on: no new Function-first run can start otherwise
Receives
Canonical working set, source units and frame
Does
Capture a bounded repeatable-read basis and admit a prepared assessment run. This is not the Map run’s analytical basis.
Produces
Sealed Function-first snapshot and admitted run
Skeletonisation choiceDeterministic unless the model skeleton is chosen
Receives
Scoped evidence and chat pointers
Does
Choose deterministic skeletonisation or the single structured model pass below. A deep run with the inverted flow on uses the other seed lane instead, unless the zero-node fallback is switched on and the deep seed returned nothing.
Produces
Skeleton nodes and channels
Free-form diagnosis seedA deep run, with the inverted flow on in the shipped configuration and able to be switched off; switched off, a deep run falls back to the seed lane on the left
Receives
Whole corpus, case frame and prior structure for orientation
Does
Generate a diagnostic seed through InvertedFlow. A content decode failure fails closed; this is not the later Diagnosis run.
Produces
Claims, actor material and diagnostic seed
Execution authorisationSeparate from preparation
Receives
Prepared sealed run
Does
An authenticated person authorises all six functions or one selected pilot function. Provider configuration alone is insufficient.
Produces
Bound execution selection and authorisation
LLM skeleton Pass 1One structured pass, only when the model skeleton is chosen; the default skips it
Receives
Scoped source evidence
Does
Produce structured node and channel proposals in one structured pass; a failure or an empty result retries, which costs more calls. Then validate deterministically, resolve source handles and write the admitted skeleton; there is no separate prose-to-formatting model pass here.
Produces
Validated, source-resolved skeleton candidates
Backward groundingWithin the inverted seed
Receives
Free-form claims and source corpus
Does
Resolve the claims back to source passages; retain the distinction between confirmed material, hypotheses and what cannot be determined.
Produces
Grounded findings and grounding resolutions
Dispatch and reservationProvider identity and budget required
Receives
Authorised selection and sealed basis
Does
Validate the selection, the provider identity and the accounting, and reserve bounded model work per turn and per tool call. Currentness is not rechecked here: whether the basis still describes the live world is settled once, at authorisation, a gate at the door rather than a thread through the run.
Produces
Admitted request or visible refusal
Factual preparation runA standard Map requested with factual preparation: deterministic skeletonisation only, then it stops
Receives
Frame and scoped evidence
Does
Wire the deterministic skeletonise stage and nothing else. Trace, audit, structuring, cross-channel work and Map reasoning are left unwired, so the run terminates after preparation. Requesting it together with a deep run is refused at the run row.
Produces
A skeleton and its captured basis, recorded as a factual-preparation run. Not a Map that Diagnosis can read.
Gap probeReads backward-grounded findings and the corpus
Receives
Grounded findings, raw corpus and prior structure
Does
Probe missing evidence. The probe itself does not emit skeleton nodes; its open questions join the inverted result.
Produces
Gap observations and open questions
Track B: function assessment loopS1 · S2 · S3 · S3* · S4 · S5
Receives
Sealed basis, compact index and selected function
Does
For each authorised function: model turn → bounded tool calls → returned passages and receipts → next turn or terminal verdict. Current product requests retain accumulated lawful exchanges within that function; a compact index does not mean all prior exchanges are pruned. Historical script overlays used a different context-pruning policy. Only selected functions run; budgets and transport failures can stop work.
Produces
Assessment verdict and retained exchanges
Seed projection and persistenceCompletes the inverted skeletonisation slot
Receives
Grounded inverted result, gaps and scoped frame
Does
Persist untyped carrier nodes and missing-element markers. Failures abandon the run. A successful run that produced no node falls back to deterministic skeletonisation only when that fallback is switched on, and it is off unless the operator turns it on.
Produces
Persisted seed nodes and missing-element markers
Quote grounding verificationChecks sealed bytes and windows actually returned
Receives
Verdict anchors and tool receipts
Does
Require byte-exact quote containment in this function’s served source windows. This proves quote provenance, not claim entailment.
Produces
Verified anchors or refused terminal/binding
Map traceBoth seed modes; conditional on observed chat signals
Receives
Skeleton and channel material
Does
Trace channel content when a chat signal is observed. Both seed modes enter the same conditional trace gate; inverted seeding does not make trace unconditional.
Produces
Channel content states, or an unrun stage
Persist assessments and findingsCan stop here; no Standard Map Audit required
Receives
Accepted terminals, bindings and anchors
Does
Persist assessments, findings and supporting receipts. Evidence status and execution status remain separate. Unassessed does not mean absent; findings remain provisional model interpretation. Standard Map Audit is not required to finish this assessment; subsequent proposal review and promotion are separate, optional work.
Produces
Durable function findings and execution outcome
Audit evidence capsulesCandidate-local assembly; joins trace states at audit
Receives
Candidate nodes and grounding events
Does
Assemble candidate-local allowed evidence for audit. Empty grounding fails closed; channel regrounding is a separate check. This is request assembly, not another universal model call.
Produces
Audit inputs and grounding closure
Derive eligible proposalsOptional feature-gated completed-run derivation
Receives
Completed assessment findings
Does
Derive immutable proposals for review when the proposal feature is enabled; derivation failure is separate from settled assessment execution.
Produces
Proposals linked to findings
Map auditNode typing and explicit abstention
Receives
Candidates, evidence capsules and trace states
Does
Audit proposed functions. The deterministic capsule gate applies to node classifications under selected evidence; channel quote regrounding has separate scope.
Produces
Typed nodes, annotations and abstentions
Human review decisionsAccept, caveat, defer, reject or withdraw
Receives
Immutable proposals and evidence
Does
Record append-only human decisions. Acceptance alone does not add canonical Map roles.
Produces
Review decision chain
Pre-structuring completeness checkAfter audit; not universally a stop gate
Receives
Typed/abstained node set
Does
Calculate audit coverage. With structuring off, this is the census location; with it on, final product census emission is deferred. Strict evaluation and product routes have different failure rules.
Produces
Coverage calculation and incomplete-audit signal
Seal reviewed reportExact proposals, decisions and currentness
Receives
Human decisions and sealed assessment basis
Does
Sign an immutable review assembly over exact identities and hashes. Signing is separate from promotion.
Produces
Signed Function-first review report
Map structuring / legacy recursion gateOn in the shipped configuration and can be switched off; the two recursion mechanisms never both run
Receives
Audited structure and focal system
Does
When wired, structuring is the recursion mechanism: ground producers, revise the structure and reuse child spawning. Otherwise the legacy recurse stage may run on its eligibility gate.
Produces
Revised nodes and eligible recursive children
Promote accepted roles into a new MapExplicit human command; not a signing side effect
Receives
Current signed report and eligible accepted proposals
Does
In the current product, recheck authority, sealed/live quotes and exact selection; append role annotations and a promotion receipt, then queue one ordinary, non-deep standard Map. With multi-role enabled, accepted effective-role nodes can be excluded from new Audit candidates; unresolved nodes, channels, child systems, structuring and reasoning can still need work. This is not a universal Audit bypass.
Produces
New Map run; no direct shortcut to Diagnosis
Final post-structuring censusWhen structuring ran; before cross-channel work
Receives
Revised typed/abstained node set
Does
Recheck and emit the final census after structuring. A coverage record does not establish variety, balance or viability, nor does every product run stop on incompleteness.
Produces
Run-scoped census and incomplete signal
Cross-channel workOn in the shipped configuration and can be switched off, and the prerequisite stages succeeded
Receives
Parent System 1 and child-system pairs
Does
Assess vertical channels after the recursion mechanism has spawned eligible children. Failed/incomplete required revision skips this stage.
Produces
Typed cross-level channel observations
Map reasoningLast analytical producer; relevant gates must pass
Receives
Revised typed structure, channels and audit abstentions
Does
Produce provisional hypotheses, gaps and actor overload observations over the revised Map. A malformed answer is retried once.
Produces
Reasoning outputs and exact reasoning inputs
Map settlement and input basisSuccessful run captures its final reasoning basis
Receives
Map structural writes and reasoning outputs
Does
Persist the Map and bind a successful eligible run to captured analytical inputs. Earlier stage outputs are written within the Map transaction; settlement is not one final universal write.
Produces
Run-scoped Map, basis and outcome
Diagnosis input captureSeparate requested run; requires an eligible current Map
Receives
Latest completed eligible standard Map and its basis; frame, evidence and observations
Does
Capture coherent repeatable-read inputs. Stale, incomplete or unavailable Map authority blocks Diagnosis; factual prep alone is insufficient. Function-first must first enter a new ordinary Map through human promotion.
Produces
Diagnosis basis bound to the Map run
Diagnosis reasoningSeparate from the inverted Map seed
Receives
Captured frame, sources, structure and observations
Does
Generate the diagnostic certificate against retained inputs.
Produces
Certificate JSON
Certificate validationRetries once on a malformed answer; on success the run continues straight into the gate
Receives
Certificate JSON and rendered evidence
Does
Validate the primary target catalogue and mechanism comparisons. Quote fields do not all share one rendered-span check; later write gates have their own scope.
Produces
Validated findings, prescriptions and gaps
Structural hypothesis reasoningOptional parallel branch when live observation keys exist
Receives
Structural seeds and grounded prompt
Does
Generate provisional prose structural hypotheses alongside the primary findings path. With no eligible observations, the path continues without this branch.
Produces
Prose hypotheses
Structural hypothesis harvestOptional branch’s formatting model call
Receives
Prose hypotheses
Does
Format the hypotheses and key them to observations. Any hypothesis left with no observation key is dropped. This does not rejoin the confidence gate, which ran before this branch started.
Produces
Structured hypotheses, appended at settlement as their own records
Diagnosis confidence gateRuns on the primary findings path, before the optional structural branch
Receives
The certificate's confirmed findings, and the pre-provider dependency capture for the Map hypotheses the prompt actually rendered
Does
Apply deterministic confidence rules, including backing-hypothesis source grounding. This does not establish causal truth.
Produces
Gated provisional findings
Diagnosis emissionRun-scoped records
Receives
Gated findings and commitments
Does
Emit prescription reasoning events while preserving finding, hypothesis and human commitment distinctions.
Produces
Retained diagnostic records
Diagnosis renderSaved records become visible
Receives
Emitted records
Does
Render retained findings with evidence and human-review surfaces.
Produces
Consultant-visible Diagnosis
Compose diagnostic reportAutomatic best-effort, or explicitly requested
Receives
Run-scoped Diagnosis findings and commitments
Does
Compose for a terminated completed/abstained prescription run only; Function-first is excluded. Composition failure does not change the analytical run outcome.
Produces
Deliverable or separate composition failure
Map inquiry agendaBranch from Map settlement, not the report
Receives
Map gaps, open needs and answered gaps
Does
Build authoritative inquiry needs for the consultant.
Produces
Questions to investigate
Function-first inquiry needsOwn post-assessment read model; not automatic chat
Receives
Valid function assessments and open needs
Does
A valid considered but not evidenced assessment can expose an evidence request on its run-review surface. An operational not assessed result stays separate and does not manufacture an absence question. This is not the Map QuestionBlend engine and does not automatically start consultant chat.
Produces
Eligible Function-first inquiry needs, separate from execution failures
Map question generationConditional generation from the Map agenda
Receives
Allowed Map inquiry needs
Does
QuestionBlend can generate on a cache miss when enabled; retained generations are distinct from a fresh question call.
Produces
Rendered or retained questions
Consultant chatUser-triggered Map inquiry branch
Receives
Agenda, available Map hypotheses and conversation
Does
Reply using the selected inquiry context. This is not a stage of every analysis run. The reply itself emits no observation events: it returns the turn to the caller and saves it.
Produces
One saved chat turn
Selected question agentsAlternative branches for selected Map question kinds
Receives
Selected question kind and permitted context
Does
Dispatch to the wired question agent appropriate to the selected question; four wired agents are alternatives, not four compulsory sequential calls. The variety-engineering agent records a grounded variety-mechanism observation per turn, dropping any it cannot resolve to a source passage.
Produces
Question-specific inquiry response, and the variety-mechanism observations
Chat extractionDurable optional job → a later Map input
Receives
Saved turn text
Does
Extract scoped chat evidence into source-backed reasoning events, including the inter-party observations a reply never emits. This does not alter a sealed Function-first run or automatically dispatch a new one.
Produces
chat_* events, the inter-party observations among them, for later Map input
Explicit audited Map resumeOptional non-deep dispatch; not a Function-first shortcut
Receives
A selected source Map run and its retained durable Audit
Does
Verify source-run identity and audit coverage with a read-only proof: no model calls or writes in that check. Only an explicit eligible non-deep continuation can skip seed, trace and fresh Audit, then enter the downstream shared Map stages. This needs its own proof; promotion does not select it automatically.
Produces
Proven continuation inputs, or refusal
Standalone InvertedFlow resultEngine stopping point; not another app route
Receives
Source corpus and case frame
Does
The standalone engine call runs diagnosis → backward grounding → gap probe, separates confirmed material from hypotheses, and returns. It does not project Map nodes or run Standard Map Audit. The integrated inverted seed above calls this engine and then projects its result into Map.
Produces
Diagnostic result without a standard Map
Answers and interview materialOptional case-level evidence contribution
Receives
Questions, human answers and interview material
Does
Record source material and queue the routing of that evidence, retaining what was admitted and the receipt for it. An answer can queue a rerun of the analysis, or record that a rerun is owed. Neither is a required step between analysis stages.
Produces
New source units, admitted evidence, routing receipt and rerun/debt record
Corrections and structural feedbackNew correction events; never rewrite old evidence
Receives
Map structure, findings and their sources
Does
Record new correction/feedback events and, when requested, select a later Map capture. This is not a return through document emission.
Produces
Correction records and later Map inputs
Analytical reviewSeparate from Function-first report signing
Receives
Findings, hypotheses and their source evidence
Does
A person records an analytical review. Human declaration, canonical promotion and report review remain distinct authorities; later Map or Diagnosis captures can consume eligible review records.
Produces
Human review records
Commitments and observed outcomesOptional follow-up; feeds later Diagnosis
Receives
Chosen actions and subsequently observed outcomes
Does
Record human commitments and later outcomes separately from the model’s prescriptions; use eligible records in a later Diagnosis capture.
Produces
Commitment/outcome records

The method

How the analysis works

The groups below describe the declared reference workflow, not a receipt for any case. The route map above shows how the paths connect; retained run records establish what actually ran.

Every stage, in readable parts 11 parts. Each stage says what it receives, what it does, what it produces, and where a person can question or change the interpretation.

The analysis is a chain of stages. Some are paid calls to a language model with a fixed instruction and a strict output contract; most are ordinary code that checks, records or renders. Every model output that makes a claim must quote the passage it rests on, and a person can question or overrule an interpretation at several points.

Prepare source material Parsing, document extraction and factual preparation of what the analysis reads (Documents, Parsing). 7 stages
  1. Parse and span documents

    Prepares data

    Turn an uploaded PDF into stored source passages with a local parser. No model is called and nothing is paid.

  2. Read a document into a ledger

    Model call

    The first, paid reading of each window of document spans (code name stage 1): the model writes a prose ledger of the entities, roles and relations the spans support, quoting them.

  3. Transcribe the ledger to structured entities and relations

    Model call

    Turn the ledger into a loose structured proposal (code name stage 2: entities, relations, quote hints) that the resolver can snap to exact spans.

  4. Recover missed relations

    Model call

    A bounded third call over the resolved entity catalogue: propose relations between known entities that the transcription missed, grounded in the same spans.

  5. Flow-trace channel signals

    Model call

    A separate model call that reads the spans for information flows with a byte-exact verbatim phrase and within-document locator, emitted as channel-signal events.

  6. Channel-signal backfill (operator script)

    Model call

    Record channel-signal observations for documents already held, so the Map and Diagnosis prompts carry provisional signals, without turning the ingest stage on.

  7. Record document evidence events

    Saves

    Write the resolved entities, relations and signals as document evidence, safe to replay after a failure through the stored extraction attempt.

Frame the case and select evidence The consultant's declared frame and the exact inputs each analysis captures (Framing, AnalysisInputs). 4 stages
  1. Declare the case frame

    Saves

    The consultant's explicit frame: system in focus, recursion level and criterion, TASCOI and the presenting problem. A human declaration, never a model output.

    Where a person steps in. The consultant declares the system in focus, the recursion level and the scope; the analysis reads that frame and may question it, never override it.

  2. Decide whether new evidence reruns the analysis

    Prepares data

    When an answer or interview lands, mint the trigger identity and fence for a rerun and track reanalysis debt, so a new run is attributable to the evidence that caused it.

  3. Select the current Map inputs

    Prepares data

    One coherent read-only snapshot of what a Map run would see now: typed nodes and channels, source units, the diagnosis packet and frame values.

  4. Capture the Diagnosis inputs

    Prepares data

    Bind the latest standard Map run and its input basis, check it is current, and assemble the render context the Diagnosis prompt is built from.

Initial analysis and gap exploration The inverted, free-form stages inside a deep Map run. 2 stages
  1. Free-form initial analysis (deep run)

    Model call

    One pass over the whole corpus by a frontier model: bold diagnosis with per-fact verbatim quotes, which the backward grounder then checks before anything is written.

  2. Gap exploration probe (deep run)

    Model call

    A normative VSM completeness sweep over the corpus and the diagnosis so far: what is absent or under-evidenced, each with a probing question. Gaps are inferences, never citations.

Build and examine the system map The structural and reasoning stages of a Map run, from the skeleton to the reasoning. 11 stages
  1. Skeleton from evidence (deterministic)

    Prepares data

    Build the first nodes and channels from extraction events without a model call. This is the default skeletonize mode.

    Where a person steps in. After the skeleton is built the consultant corrects the working set: merges duplicate actors, renames, retypes; each correction is a recorded event the next run reads.

  2. Skeleton by model

    Model call

    Identity-only proposal of nodes and channels from source units and chat pointers, with verbatim quotes and no S-labels.

  3. Trace signal flow

    Model call

    Follow each observed signal through the skeleton: origin, first visible point, degradation state and any algedonic bypass, quoting the signal text verbatim.

  4. Assemble the audit evidence capsule

    Prepares data

    Per audit candidate, select the CLOSED set of evidence the audit may see — the canonical grounding chain, one extraction-relation endpoint hop, and the source units those groundings cite. The audit request is assembled from this capsule, never from the case corpus.

  5. Audit: classify or abstain

    Model call

    The canonical typing authority: assign S-labels to nodes and Beer channel sub-types to channels only where the evidence entails it, otherwise abstain with a reason.

    The gate exists to avoid false holism: a map that looks complete because the model filled the empty boxes.

    Where a person steps in. Audit classifications and abstentions are visible per node; a consultant's correction supersedes the model's classification and is kept as its own record.

    What this stage will not admit

    These are the gate's own refusal sentences, quoted from its declared vocabulary. They are not a record of declines on any case.

    • Semantic evidence gate declined the typing because a name or role title is not evidence of organisational function.
    • Semantic evidence gate declined the typing because the target-bound quote supports the proposed function but does not distinguish a formal mandate from described activity without overclaiming enactment.
    • Semantic evidence gate declined the typing because the target-bound evidence negated or prohibited the proposed function rather than assigning or describing it.
    • Semantic evidence gate declined the typing because the same target-bound evidence supports more than one organisational function and the scalar label cannot preserve that role overload honestly.
    • Semantic evidence gate declined the synthesis typing because a contribution did not verbatim-anchor to an exact grounding event in the cited source.

    The gate has 28 declared grounds; these are 5 of them.

  6. Completeness census

    Validates

    After audit: every node must be typed or explicitly abstained. On the product route an incomplete census is recorded and structuring, recursion and reasoning still run; only the strict evaluation routes stop on it.

  7. Structuring: System One value producers

    Model call

    Identify the System One value producers from what the system actually does (POSIWID), correct environment/S1 confusions as events, and feed the recursion machinery.

  8. Cross-level channels

    Model call

    For each parent S1 / child VSM pair, classify the five vertical strands (resources bargaining, accountability, legal/corporate, audit, algedonic) strictly from verbatim evidence.

  9. Reason as the VSM: hypotheses

    Model call

    The heaviest Map pass: over the typed structure and the diagnosis packet, emit structural, systemic and general hypotheses, S3/S4 balance, autonomy and gaps, each verbatim-grounded.

    Where a person steps in. Each hypothesis shows its quotation; a consultant can record an analytical review that challenges, confirms or rewrites it, with the review kept beside the original.

    **Naming a mechanism.** A delay, a bottleneck, a single point of failure, an overload, a failed escalation — each of these is a claim about how the organisation BEHAVES, not a description of how it is ARRANGED. Name one only when a quoted passage states it — someone reports the wait, the queue, the handover that did not happen, the decision that could not be made without one person — or when several quoted observations together support it, and then the hypothesis must say which observations contribute, what remains uncertain, and
    from the Map reasoning prompt
  10. Recurse: spawn child systems

    Model call

    Per typed S1 candidate, decide whether to spawn a child VSM (own operations and own boundary, at least two source units, verbatim criterion quote) or abstain.

  11. Canonical structure and reasoning events

    Saves

    Every Map output lands as events through the append chokepoint with a substring tripwire on quotes; canonical nodes and channels are projections of those events.

Assess VSM functions The separate Function-first assessment route. 3 stages
  1. Function-first assessment loop (Track B)

    Model call

    A tool-using loop over a frozen basis: for each of six VSM functions the model searches evidence, opens dossiers, fetches passages and records a verdict with quotes.

  2. Quote grounding verification

    Validates

    Every quote in a verdict must be byte-exact in a source the loop actually opened.

  3. Function assessment findings

    Saves

    Findings, turns, tool receipts and transport receipts retained per run.

Develop the diagnosis The prescription certificate, structural hypotheses and the report (Diagnose, PrescriptionOps). 8 stages
  1. Diagnosis certificate

    Model call

    The diagnosis itself: confirmed findings, gap probes, prescriptions and structural inferences with critical flow, breakpoints and sway vectors, in plain language, against a closed target catalogue. Separation-of-duties overlaps are never confirmed.

  2. Certificate and citation validation

    Validates

    Validate the certificate against the closed target catalogue, resolve mechanism comparisons, and check every quote against the rendered source spans.

  3. Structural hypotheses: reasoning

    Model call

    Free-prose, forward-looking hypotheses anchored to established observation keys (quote-exempt Tier 3), extending the certificate's cross-cutting inferences.

  4. Structural hypotheses: harvest to JSON

    Formats

    A formatting-class model turns the prose into hypothesis rows keyed to observation keys, predicted risk and how to verify; keys outside the live set are dropped.

  5. Confidence gate

    Validates

    A finding with no source-unit grounding is downgraded to a gap probe; confidence is capped at its backing's ceiling; a failed dependency lookup fails closed.

  6. Translate and record findings

    Saves

    Translate the gated certificate into prescription_* reasoning events and prescription commitments, through the strict substring tripwire.

    Where a person steps in. Findings are proposals until a person commits, revises or withdraws them; the case record shows which state each is in.

  7. Generate the diagnosis report

    Formats

    Compose the deliverable from one Diagnosis run's retained commitments and findings. A separate action from the Diagnosis itself and from reviewing it; no model call.

  8. Diagnosis page

    Renders

    Present the retained findings, their sources and history; nothing is generated here.

Prepare investigation questions The routes that produce questions: the deterministic agenda and the model-rendered engine. 3 stages
  1. Question agenda (deterministic)

    Prepares data

    The authoritative agenda: Map gaps, open inquiry needs and answered gaps, ranked by the server-side inquiry policy without a model.

  2. Question rendering by model

    Model call

    Render the server-selected inquiry needs as plain-language questions with input types and grounded options; the model may not reword the selected question.

    Where a person steps in. Investigation questions are answered by the consultant in the conversation; an answer becomes a cited source unit the next analysis can read.

  3. Differential questions (Engine C)

    Model call

    Episodic differential-diagnosis questions that split competing hypotheses.

Conversation The consultant conversation: reply, turn extraction, side agents (Chat). 3 stages
  1. Conversation reply

    Model call

    The synchronous next move in the consultant conversation: the agenda owns question selection; surfaced quotes are substring-checked before the reply is returned.

  2. Turn extraction

    Model call

    Asynchronously read each consultant turn into typed reasoning events with a verbatim quote and confidence; no VSM typing, no hypotheses.

  3. Per-question side agents

    Model call

    Side-thread drafting agents for one selected question, sharing a prefix prompt plus a kind-specific prompt; confirmed-only emissions with a verbatim quote.

Human feedback Answers, interviews, reviews, promotions, corrections and outcomes recorded by people. 9 stages
  1. Consultant answers a question

    Saves

    A consultant's answer becomes a cited source passage and closes the gap it answers; a later correction is a new event, never an edit.

    Where a person steps in. An answer to a gap question is recorded with its author and time and is cited like any other source.

  2. Map corrections and identity confirmations

    Saves

    Human corrections to the Map and to observations, and confirmed actor identities, recorded as new events that the next run reads.

  3. Duplicate-actor suggestion

    Model call

    Asks whether two canonical records denote the same real-world entity, given the passages cited for each and any passage naming both. A SUGGESTION beside the pair: it merges nothing, and the consultant's own ruling remains the only write to structure.

    Where a person steps in. A model may suggest that two records are the same actor; only the consultant's ruling merges them, and the ruling names who decided and why.

  4. Staff interview answers

    Saves

    Observed staff testimony captured against the interview agenda becomes cited source passages, read by the Map and the Diagnosis like consultant answers.

  5. Route human answers into admitted evidence

    Saves

    A durable job that takes consultant, staff-interview and chat-turn answers and records their analytical evidence admission with a routing receipt, so later analyses read admitted evidence rather than raw text.

  6. Intervention outcomes and failed attempts

    Saves

    After a Diagnosis: decisions on prescriptions, observed interventions, judged outcomes, learning, and recorded failed attempts, which the next Diagnosis reads.

  7. Consultant analytical review of a finding

    Saves

    The consultant records a stance on one named finding, a gap probe, a critical flow, a breakpoint, a structural hypothesis or a Map gap, with cited source passages. It is one recorded review, append-only: a later review or withdrawal retires it.

  8. Review, seal and promote assessed structure

    Saves

    Human decisions on function proposals, sealing the assessment report, and promoting accepted roles into the structure the next Map reads. Separate from reviewing a Diagnosis finding and from generating the report.

    Where a person steps in. A sealed function report can be promoted into the structure only by a person; the promotion is an event with its own authority record.

  9. Contextual contribution

    Saves

    Neutral historical context recorded on the framing with exact prompt, response, speaker and capture time; not a source unit, not an answer, excluded by analytical consumers.

Evaluation Opt-in probes and harnesses; not a product path. 2 stages
  1. Corpus harness (Wells end-to-end)

    Prepares data

    The parity acid test: creates an organisation and framing, ingests a corpus through the production chat-turn, extraction and embedding entrypoints, and runs Map and Diagnosis, to score structural-typing parity.

  2. Live LLM evaluation probes

    Model call

    Grounding and S4-discrimination probes that reuse the audit system prompt with a synthetic frame.

Shared machinery Run lifecycle, transport, structured-output validation, embeddings. 4 stages
  1. Run lifecycle: claim, heartbeat, terminate

    Saves

    Every Map, Diagnosis and function-first run is an AnalysisRun row claimed for dispatch, heartbeated while it works, guarded against a newer superseding run, and terminated once with its outcome and input basis.

  2. Crash sweep

    Saves

    A scheduled table scan that marks runs whose heartbeat stopped as crashed.

  3. Structured output: validate and retry

    Validates

    The one loop every schema-bound call shares: truncation guard, JSON repair, the caller's validator, and one retry on a malformed answer.

  4. Embeddings

    Model call

    Vectorise events and canonical nodes for recall; no prompt, an embedding request.

The complete workflow diagram Every stage and every declared connection, drawn from the same declaration the product tests against. Conditional routes are drawn because they exist, not because they all ran on any particular case. This is the software's workflow, not the organisation's VSM map.
PREPARE SOURCE MATERIAL FRAME THE CASE AND SELECT EVIDENCE INITIAL ANALYSIS AND GAP EXPLORATION BUILD AND EXAMINE THE SYSTEM MAP ASSESS VSM FUNCTIONS DEVELOP THE DIAGNOSIS PREPARE INVESTIGATION QUESTIONS CONVERSATION HUMAN FEEDBACK Parse and span documents Read a document into a ledger A Transcribe the ledger to structured entities and relations Recover missed relations Flow-trace channel signals A Channel-signal backfill (operator script) Record document evidence events Declare the case frame Decide whether new evidence reruns the analysis C Select the current Map inputs Capture the Diagnosis inputs Free-form initial analysis (deep run) A Gap exploration probe (deep run) D Skeleton from evidence (deterministic) Skeleton by model A Trace signal flow A Assemble the audit evidence capsule B Audit: classify or abstain B D Completeness census B D Structuring: System One value producers B Cross-level channels A B Reason as the VSM: hypotheses A Recurse: spawn child systems A B Canonical structure and reasoning events A Function-first assessment loop (Track B) A Quote grounding verification A Function assessment findings Diagnosis certificate D Certificate and citation validation A Structural hypotheses: reasoning Structural hypotheses: harvest to JSON Confidence gate A D Translate and record findings A Generate the diagnosis report Diagnosis page Question agenda (deterministic) D Question rendering by model D Differential questions (Engine C) Conversation reply A D Turn extraction A Per-question side agents A Consultant answers a question C Staff interview answers C Route human answers into admitted evidence C Map corrections and identity confirmations C Duplicate-actor suggestion C Intervention outcomes and failed attempts C Consultant analytical review of a finding C Review, seal and promote assessed structure C Contextual contribution
Model call (a paid request to a language model)
Ordinary code, a check or a saved record
Conditional stage (runs only on some routes)
Colour band: which part of the journey the column belongs to
Data passed to the next stage, on the ordinary consulting route
The same, on a side route (deep run, function-first, chat, feedback)
Branch: taken only when its condition holds
Retry
A person's later contribution flowing back into the analysis

What the talk claims, and where it runs

A Every claim the model promotes quotes the passage it rests onDeclared exception: structural hypotheses are keyed to recorded observations rather than to a quotation.
B Structure is linked at the next recursion level only when an evidence gate passes
C A person's corrections persist through reanalysis as their own recorded eventsEight of the nine feedback stages plus reanalysis; contextual contributions are deliberately kept out of analysis.
D Unresolved gaps become diagnostic questions rather than invented answers

A stage is marked only where the catalogue entry itself says it does that work; a stage can carry more than one claim.

Vocabulary Terminology, with sources

The vocabulary on this page is Beer's, with two borrowings that are not his. Each entry says how the software uses the term, what the source actually says, and where the two differ. Where they differ, the source wins and the software is the thing that needs work.

Variety a measure of how many distinguishable situations something has to deal with.
The source says
Beer: "a measure of complexity: the number of possible states of a system" (Diagnosing the System for Organizations, 1985, p. 35).
Where ours differs
Ashby, whose term it is, warns that the count is not a property of the thing. "a set's variety is not an intrinsic property of the set: the observer and his powers of discrimination may have to be specified if the variety is to be well defined" (An Introduction to Cybernetics, 1956, S.7/6). The system does not measure variety and makes no variety claim about any organisation.
Requisite variety the reason a regulator needs a repertoire matched to what it has to absorb.
The source says
Beer: "only variety can absorb variety", which he attributes in the same line to Ashby's Law (Diagnosing the System for Organizations, 1985, p. 35). Ashby's own statement: "only variety in R can force down the variety due to D; variety can destroy variety" (An Introduction to Cybernetics, 1956, S.11/5).
Where ours differs
Ashby counts the regulator's moves, not the products of the system being regulated. Counting what an organisation can produce is not a measure of requisite variety, and that gloss is mine, not his.
Recursion, system in focus, criterion and level the map has a system in focus, a level above it and levels below, and a child system is linked only when the evidence gate passes.
The source says
Beer: recursion is "a next level that contains all the levels below it" (Diagnosing the System for Organizations, 1985, p. 16), and "the System-in-focus is in the centre of a higher level of recursion, in which it is embedded, and it contains a set of viable systems which exist at the next lower level of recursion" (same book).
Where ours differs
the level is not the same thing as the criterion you unfolded along. Pérez-Ríos: "we could take recursion criteria three or recursion criteria one, and then we will get a different pipe, a different pipe... But the organisation is one, is the system in focus" (Metaphorum webinar, 11 May 2022, at 15:18). The software records one criterion at a time, declared by a person. It cannot tell you that a different criterion would have been better.
System One the value producers, the parts that do what the organisation is for.
The source says
Beer: "The set of these embedments will be known as SYSTEM ONE of the System-in-focus" (Diagnosing the System for Organizations, 1985, p. 19).
Where ours differs
Beer's test is stricter than ours. For Beer an element of System One is itself a viable system embedded in the system in focus, not merely the part that does the work. "Value producer" is my shorthand and the phrase is not in Beer. The software's floor asks whether a carrier performs the primary service or operation. It does not test viability, so it can type something Beer would refuse.
System Two coordination between the parts, and the only thing an S2 finding is allowed to mean.
The source says
Beer: "Its function is not to command, but to damp oscillations" (Diagnosing the System for Organizations, 1985, p. 68).
Where ours differs
Beer warns in both directions. System Two is "almost totally misunderstood and under-represented in contemporary management technique" (same, p. 66), and there is "a compulsion on users of the VSM to cram all sorts of corporate activities into System Two" (same, p. 70). He is explicit that "accountability does not reside in System Two" (same book), so an S2 finding here is never an accountability finding.
System Three the here and now, resource allocation across the parts, performance management.
The source says
Beer: "System Three, then, is responsible for the internal and immediate functions of the enterprise: its 'here-and-now', day-to-day management" (Diagnosing the System for Organizations, 1985, p. 86), with the resource bargain as "the 'deal' by which some degree of autonomy is agreed between the Senior Management and its junior counterparts" (same, p. 38).
Where ours differs
for Beer, accountability upward is a variety attenuator, not a reporting line. "in principle the homeostatic message upward needs to be only 'OK'" (same, p. 81). More reporting is his pathology, not his remedy, and the software has no representation of that at all.
System Three-star, and our Audit stage evidence that someone reaches past the management line into the operations themselves.
The source says
Beer: "Such mechanisms work sporadically" and "penetrate straight to the operations themselves", and "These procedures, which may generically be called 'audits', are indicated as the sixth vertical channel" (Diagnosing the System for Organizations, 1985, p. 82). He adds that "routine and regular audits surrender a large part of the variety they generate to no purpose whatsoever" (same, p. 85).
Where ours differs
this one is a name collision and I would rather disclose it than rename it quietly. The pipeline has a stage called Audit. It is a typing pass the software runs over the whole map on every run. Beer's System Three-star is sporadic, high-variety, and an organ of the organisation, not of the consultant's instrument. Ours is routine by design, so it cannot be the thing Beer means. Beer also flags this label and System Two as the likeliest places to go wrong.
System Four outside and future scanning, strategy and intelligence.
The source says
Beer: the phrase is his, "the outside-and-then", and he asks more of it. "SYSTEM FOUR is not only concerned to manage the outside-and-then, but to provide self-awareness for the System-in-focus" (Diagnosing the System for Organizations, 1985, p. 115).
Where ours differs
we type the scanning face. We do not evidence the organisation's model of itself. That is a gap, not a simplification.
System Five policy, identity, the ground rules.
The source says
Beer: System Five is there "to supply logical closure to the viable system", and closure means "self-reference: the assertion of identity" (Diagnosing the System for Organizations, 1985, p. 129).
Where ours differs
Beer also gives System Five the job of monitoring the Three-Four homeostat, which he calls "the organ of ADAPTATION for the enterprise" (same, p. 120). The software does not assess that monitoring relationship.
Homeostat a paired relationship whose balance can be assessed, and at present we assess one of them, between Three and Four.
The source says
Beer defines the state rather than the device: "stability of a system's internal environment despite the system's having to cope with an unpredictable external environment" (Diagnosing the System for Organizations, 1985, p. 16).
Where ours differs
Beer counts thousands of them. Every line between two dotted points on his charts is a homeostatic loop, "each being susceptible to cybernetic analysis" (same, p. 146). We look at one. "Homeostat assessment" should not be read as though we had looked at the rest.
Amplifier, attenuator, transducer three mechanisms a channel can carry, recorded per direction, with evidence attached to each.
The source says
Beer: an attenuator is "a device that reduces variety", an amplifier "a device that increases variety" (Diagnosing the System for Organizations, 1985, p. 35). Transduction is his Third Principle of Organization: "Wherever the information carried on a channel capable of distinguishing a given variety crosses a boundary, it undergoes transduction; the variety of the transducer must be at least equivalent to the variety of the channel" (same, p. 47).
Where ours differs
the three primitives are Beer's. Arranging them as six cells, three mechanisms in each direction of a channel, is a representation choice taken from Clemson's 1994 figure so that the evidence has somewhere to go. Beer counts transducers by boundary crossings, not by direction cells.
TASCOI part of the framing a consultant declares before the analysis runs, alongside the system in focus, the recursion level and the presenting problem. It is a human declaration and never a model output.
The source says
Espejo and Reyes: "the mnemonic TASCOI where T stands for the canonical form of the Transformation, A stands for the Actors performing the transformation, S stands for the Suppliers, C stands for the Customers, O stands for the Owners and I stands for the Interveners" (Organizational Systems: Managing Complexity with the Viable System Model, 2011, p. 126).
Where ours differs
TASCOI is Espejo's, not Beer's. Their "Owners" is not the everyday reading either. It means "those persons or bodies in the organization that have an overview of the transformation and have the responsibility for adjusting performance to meet some criteria of effectiveness" (same, p. 126). In the software, a framing declaration is consultant background. It is never treated as evidence.

Practitioner pilots

I am looking for a few practitioners to try one bounded inquiry after the conference: a real question from an organisation you already work with, appropriate permission to use the material, and a return visit after the first analysis.

Richard Gunther, richard@recursive.systems

Research direction: Agents

An agent could help use the application: retrieve a case, inspect an argument or propose a question. Internal agents can be connected through application operations; external assistants through authenticated tools.

A team of agents could itself become the subject of inquiry. What work must continue? Who sets its purpose? What can each worker decide, and what feedback reaches someone able to change the arrangement?

That connects the work to AI governance and, potentially, research on AGI safety. The question is how a proposed interpretation becomes permitted action, and how people can challenge or revise it before an error propagates. The application here would be an inspectable inquiry into purpose, authority and feedback, not a claim that this prototype has solved alignment.

Research direction. Nothing in this section is built.