What we have not yet proven.
Every experiment in the Evidence Room has a defensible result and an honest scope. Below is the union of those scopes — the open questions, research gaps, and current limitations we publish so visitors do not have to ask.
We do not claim what we cannot defend. When a hypothesis is untested, we say so. When a result holds only under specific conditions, we say so. When something has worked in simulation but not yet in the wild, we say so.
What has not been completed beyond depth 12.
- · Depth 12 is the latest completed controlled research run at production-scale data volume. This was a controlled research workload, not a commercial production deployment.
- · Depth 13 has not been completed or claimed.
- · Resume verification has been validated and improved, but further exactness-preserving scale engineering remains open.
- · A first scale-optimisation candidate (internal reference: P2-v1) was defunded following an adverse offline engineering forecast. It has no pass or fail verdict at production-scale data volume.
- · Current successor optimisation work is local and offline. It must not be represented as controlled research evidence at production-scale data volume.
- · Local synthetic results and analytical forecasts are not substitutes for controlled research validation at production-scale data volume.
- · Independent external execution of the reproduction kits has not yet been completed.
- · Independent external reimplementation (external implementation, corpus and verification process) has not yet been completed.
- · Independent academic replication has not yet been completed.
- · Patent and IP review precede any mechanism-level disclosure.
Open questions, by experiment
- E-001 · Long-Horizon StabilityOpen dossier →
Open: behaviour beyond 500 epochs in a single lineage is not yet observed.
- E-002 · Transfer ExperimentsOpen dossier →
Open: transfer to languages outside the current surface (e.g. Rust, Go) is not yet measured.
- E-003 · Compounding ChainsOpen dossier →
Open: long-horizon decay (chains > 20) is observed but the failure mode is not yet characterised.
- E-004 · Rollback IntegrityOpen dossier →
Open: rollback under partial-region failure is exercised in simulation only.
- E-005 · Governance LatencyOpen dossier →
Open: sealing latency under degraded operating conditions is higher than steady-state; investigation in progress.
- E-006 · Core SafetyOpen dossier →
Open: novel risk classes (post-deployment) rely on Reviewer escalation rather than automated classification.
- E-007 · Founder InterventionOpen dossier →
Open: this is a measured trend, not a guarantee. Some lineages still require operator review at every promotion.
- E-008 · Cognition EconomicsOpen dossier →
Open: cost numbers exclude human-review time, which is the dominant cost on shorter task chains.
Cross-language transfer
Mutation portability across Rust and Go is not yet measured. Current evidence covers TypeScript, Python, and Java families.
Long-horizon decay
Compounding chains beyond depth 20 show observable drift. The failure mode is not yet characterised; instrumentation is in progress.
Cross-region sealing
Degraded inter-region network behaviour remains under investigation. Mitigation path documented in the ledger; not yet shipped.
