Historical scope and interpretation
Every row except MCP-Mark is a recorded Qwen3-8B B4 composition-corpus run from docs/results.jsonl (2026-07-12..15); B4 used 432 composition worlds. The MCP-Mark row is NOT B4: docs/results.jsonl line 98 records it as the gate_scale adapter (narrow gym worlds, standard suite, fs+pg, k=4), and B4 was never evaluated on MCP-Mark — line 106 states that every prior external null (tau2 parity, BFCL flat, MCP-Mark flat) was measured on the gate_scale lineage and none of them tested composition. Within-row base/tuned comparisons use the same stated harness; cross-row absolutes are not treated as interchangeable, and the MCP-Mark row must not be read as evidence about B4. Qwen3-14B remains null because no 14B training/evaluation run is recorded. The current customer-created /sandbox world has executable validation but has not itself been isolated as the causal training corpus for these results.
τ²-bench retail (Qwen3-8B B4 · paired n=78): Recorded paired B4 result: 95% CI [+6.8,+23.8]pp, p=0.0004. The broader n=100x4 run is ledgered; the paired join contains 78 tasks.
τ²-bench telecom (Qwen3-8B B4 · n=100): Recorded never-trained-domain result with deterministic ENV_ASSERTION reward, p<1e-8. The tuned arm completed 181/200 planned episodes.
τ²-bench airline (Qwen3-8B B4 · n=50): Recorded paired unsteered result; sign-test p=0.327. This is an honest non-significant boundary, not evidence of airline lift.
τ²-bench three-domain average (Qwen3-8B B4): Derived from the recorded retail, telecom, and airline domain results. The base nearly matches the paper's 26.2, but 41.3 remains well below the paper's 61.8 target.
BFCL V4 multi-turn (Qwen3-8B B4 · all 800): Recorded full multi-turn slice using the official data and checker; tuned-only 192 vs base-only 33, McNemar p<0.0001. This is a slice result, not the paper's full BFCL-V4 aggregate.
BFCL V4 AST + abstention (Qwen3-8B B4 · n=696): Recorded same-session comparison with the corrected decoder and identical vLLM harness for both arms.
MCP-Mark filesystem + postgres (Qwen3-8B · n=204): Recorded gate_scale-adapter result (NOT the B4 composition corpus, which was never evaluated on MCP-Mark): postgres was exactly flat (6.0%→6.0%); the combined one-point decrease is floor-level noise across 204 tasks (Fisher p=0.79), not a regression. No MCP-Mark improvement was demonstrated, and none was refuted for B4 — B4 has no MCP-Mark measurement.
Qwen3-14B composite-world benchmark battery: Not run. Configuration and benchmark locks exist, but no recorded 14B checkpoint or evaluation result exists in this repository.