Hadi SoufiAI Systems Architect

Research / Paper 05

SSRN Working Paper · posted 29 September 2026

Order Invariance at the Last Language-Model Boundary

A Pre-Registered Test of an LLM Evidence Synthesizer Inside a Deterministic Trading Architecture

Hadi Soufi

Abstract

Large language models are known to be sensitive to the order in which they read their inputs. How much that sensitivity matters once a model sits inside a trading system whose final decision is made by deterministic code is not known. This paper reports a pre-registered test at the only point in a production multi-agent trading architecture (ZVAKTHOR) where a language model reads an ordered list: a Research agent that synthesizes five agent votes into one directional vote. Using the production prompt, 24 synthetic evidence sets in three strata, and 3,840 calls to deepseek-flash at temperature 0, we compare all 120 orderings of each set against 40 repeats of one fixed ordering, so that order effects are measured against the model's own noise. The mean diversity increase under permutation is 0.028 (Holm-adjusted p = 0.042 within the amended two-test family; not significant under the originally registered four-test family). The point estimate is below the pre-registered smallest effect size of interest (0.05), but effects of that size cannot be excluded, and equivalence is not shown; the registered verdict is inconclusive. We find no evidence of a primacy or recency effect (p = 0.49). Intrinsic noise is large where evidence conflicts: under an identical input, the modal direction in the balanced-conflict stratum is returned in only 72% of repeats. Traced through a replica of the production aggregator, no change in the synthesizer's output direction altered any consensus result: by construction of the stimuli and the rule, consensus is unreachable in 15 of the 16 conflict sets whatever the synthesizer outputs, and all 43 observed changes arose from confidence variation in a single near-threshold set. At this boundary the practical risk is nondeterminism rather than order, and how much of it reaches a decision depends on the aggregation rule, the weight it assigns, and where inputs fall relative to its threshold.

Read on SSRN

Pre-registration: OSF Registries, DOI 10.17605/OSF.IO/M398V

Code and data: zvakthor-order-sensitivity-2026

Cite

Soufi, H. (2026). Order Invariance at the Last Language-Model Boundary — A Pre-Registered Test of an LLM Evidence Synthesizer Inside a Deterministic Trading Architecture. SSRN Working Paper. https://doi.org/10.2139/ssrn.7539442

@misc{soufi2026order,
  author       = {Soufi, Hadi},
  title        = {Order Invariance at the Last Language-Model Boundary — A Pre-Registered Test of an LLM Evidence Synthesizer Inside a Deterministic Trading Architecture},
  year         = {2026},
  howpublished = {SSRN Working Paper},
  doi          = {10.2139/ssrn.7539442},
  url          = {https://doi.org/10.2139/ssrn.7539442}
}