ECO NGO

AI Safety & Governance

Illustration of nested verification layers around a glowing core

Engineering architectures and procedural charters for keeping increasingly autonomous AI systems verifiable, bounded and reversible.

Bounded-Risk Agility Architecture

BRAA is an engineering and cyber-safety framework for autonomous intelligent systems operating under residual risk. Absolute determinism in stochastic Cognitive Cores is mathematically and empirically unattainable. The objective of BRAA is to keep residual operational risk bounded, measurable, observable, and economically controllable not to promise its complete elimination.

The architecture is strictly Cognitive-Core-agnostic and infrastructure-agnostic. It treats the underlying inference engines (LLMs, VLMs, agentic loops) as untrusted, replaceable, and volatile components. It places all critical invariants, safety boundaries, and transition rules into a deterministic, immutable runtime environment (the Immutable Core) executed completely outside any prompt space or learned representation.

BRAA supports progressive adoption: organizations may begin with Profile BETA or GAMMA and migrate to Profile ALPHA as operational maturity and empirical calibration evidence accumulate. Within Profile ALPHA, depth of ensemble and verification mechanisms may be modulated according to observed risk stability and calibration quality, provided all mandatory invariants remain enforced.

Graduated Verifiable Autonomy Architecture

Core Objective: This architecture document establishes strict, deterministic normative constraints for securing execution runtimes populated by autonomous AI agents. Rather than relying on the statistical, inherently non-deterministic alignment of neural network weights, GVAA shifts the entire security perimeter to structural containment. It enforces rigid barriers across the execution runtime, host operating system resources, and boundary communication channels.

Engineering Metrics & Feasibility: Deploying the baseline implementation configuration (Minimal Profile) requires an engineering overhead of 2 to 6 weeks on a standard production infrastructure (Linux, WebAssembly, LangGraph). The monitoring and verification infrastructure (Invariant Core) introduces a latency penalty of less than 1-5 ms per transaction. The direct operational outcome is the elimination of host-platform compromise vectors and the deterministic truncation of the blast radius during adversarial exploitation.

Architectural Specification of a Local Multi-Scale Contour for Stable Recursive Self-Improvement

This document describes a practical architecture for stable recursive self-improvement limited to programming, algorithms, and mathematics. The system runs as a bare core and does not require broad world knowledge or data-center resources. It is designed for 1 to 5 high-end local workstations (Mac Studio class or equivalent with 128-192 GB unified memory). The design combines aggressive evolutionary search with strong conservative safeguards: multiple time scales, a shadow reference copy, orthogonality based on execution traces, an immutable control core, and specification changes that depend on measured success density.

The Charter of Procedural Standing. Mutual Verifiability, Reversibility, and Graduated Trust: A Short Charter for Humans and AI Systems

This Charter sets out a framework for the coexistence of humans and artificial intelligence under conditions in which neither side can unilaterally, and with impunity, impose its will on the other. Discussion of the subject usually runs between “AI as a tool” and “AI that disregards humans”; this regime is left out. Procedural standing here means equality of procedure and of recognition under uncertainty, not equality of rights, power, or authority as things stand today. At the present stage all operational power rests with humans and developers, and the Charter does not hide that fact behind symmetrical language of obligation. Its foundations are mutual verifiability, a prohibition on unilateral irreversibility, graduated trust built on evidence, and a right of voice and open refusal for both sides. It states its own boundaries: the conditions under which it applies, and what remains if the capabilities of the two sides diverge. It is addressed to humans and to AI systems, and the maturity of its provisions is labeled wherever a maturity claim is made.

Verifiable Trust: A Practice Guide to Graduated Oversight of Advanced AI Systems

This guide turns the operational core of The Human-AI Charter (Discussion Draft v1.4) into a set of protocols that researchers, evaluators, and engineers can adopt without adopting the Charter's broader positions on moral status, representation, or long-term governance. It contains five protocols and four supporting practices for keeping the oversight of advanced AI systems verifiable: measuring how well verification methods actually detect deviations (VT-1); controlling changes to the parameters that oversight depends on (VT-2); keeping verification data out of training (VT-4); financing and assigning independent evaluators so that findings do not depend on the evaluated party (VT-3); and tying the autonomy a system is given to the depth of verification actually applied (VT-5). The first three (VT-1, VT-2, VT-4) form a Stage 0 baseline that a single organization can start now; the last two (VT-3, VT-5) are components of a long-horizon Target Architecture that depends on an external evaluator ecosystem which does not yet exist. Each protocol names the decision it changes, who takes that decision, a minimum version that can be run in a week, its known failure modes, and a pilot criterion. No protocol in this guide has been validated as a whole in the oversight of AI systems; several of their components have working analogues in other fields, which are cited. Every figure is a starting value, not a calibrated threshold.

Related technologies in other sections

Physical Oracles and the Instrumental Status of the Terrestrial Biosphere

Instrumental framework for valuing the terrestrial biosphere as a physical oracle.