ainfra · Documentation v1.0.0-alpha.7
ainfra · v1.0.0-alpha.7

Documentation

Printed August 22, 2026 · 15 pages

Contents

  1. Introduction
  2. Quick Start
  3. Installation
  4. Configuration
  5. Output, inventory, and Ansible
  6. Hetzner private K3s template
  7. Usage
  8. Guarded MCP server
  9. Templates
  10. AI template-authoring guide
  11. Reviewed OpenTofu Plans
  12. References
  13. Tutorials
  14. Contributing
  15. Roadmap

Introduction

ainfra turns reusable infrastructure templates into repeatable, reviewable deployments. It validates explicit contracts, locks template content, and coordinates established provisioning and configuration tools.

Templates combine OpenTofu for infrastructure provisioning with Ansible for system configuration. ainfra verifies the boundary between lifecycle steps and retains sanitized evidence without hiding either underlying tool.

The Go-based v1 line is in alpha. Phases 0 through 9 are released as v1.0.0-alpha.9, including the local template-authoring conformance contract, clean-room proof, and the first provider-backed production candidate. The Hetzner private K3s baseline passed its cost-approved disposable live certification, tunnel-only management check, and independent teardown audit.

Quick Start

The current alpha CLI can validate a local deployment, lock a local or Git template source, create a reviewed OpenTofu plan, and apply only that exact saved plan.

Install the latest release:

sh
curl -fsSL https://raw.githubusercontent.com/projectious-work/ainfra/v1.x-release/scripts/install.sh | bash

If ~/.local/bin is not already on PATH, add it before continuing. See Installation for version pinning and signature verification.

Start by checking local prerequisites:

sh
ainfra doctor environment

Create a minimal deployment directory when starting from scratch:

sh
ainfra init example-deployment

This writes example-deployment/ainfra.yaml and an idempotent marked .gitignore entry for local .ainfra/ evidence. It does not create secrets, backend resources, or infrastructure. Replace the placeholder local template reference with the source you intend to lock.

Resolve and materialize the source, then create its canonical lock:

sh
ainfra template lock example-deployment

For Git sources this is the only step that resolves the requested mutable ref. It records the immutable commit and verified content digest. For intentional source, ref, or content changes, use the explicit update path:

sh
ainfra template update example-deployment

Then diagnose a deployment directory or its manifest:

sh
ainfra doctor deployment path/to/deployment
ainfra doctor path/to/deployment --format json

For a checked-out template source, run:

sh
ainfra doctor template path/to/template

Doctor is read-only unless --reconcile is explicitly supplied and its plan is confirmed. It verifies the lock, contained cache entry, digest, and local source drift without reacquiring mutable Git refs.

Create a reviewed apply plan:

sh
ainfra plan path/to/deployment

Review the structural action counts and digests in the result. Then copy its exact next command, for example:

sh
ainfra apply path/to/deployment \
  --plan 20260814T120000Z-0123456789abcdef0123456789abcdef

Immediately before OpenTofu starts, ainfra reverifies the deployment, native input files, template binding, workspace, OpenTofu executable and version, and saved-plan bytes. On v1.x-dev, continue with the Phase 5 output, inventory, and Ansible workflow.

Installation

ainfra publishes archives for Linux and macOS on amd64 and arm64. The official installer detects the current platform, downloads the matching release, verifies its published checksum, and installs the binary:

sh
curl -fsSL https://raw.githubusercontent.com/projectious-work/ainfra/v1.x-release/scripts/install.sh | bash

The default destination is ~/.local/bin. Ensure that directory is on PATH, or select a version and destination explicitly:

sh
curl -fsSL https://raw.githubusercontent.com/projectious-work/ainfra/v1.x-release/scripts/install.sh \
  | VERSION=1.0.0-alpha.7 INSTALL_DIR=/usr/local/bin bash

Every release includes checksums.sha256 and its checksums.sha256.sigstore.json signature bundle. Checksum verification is mandatory. When Cosign is available, the installer also verifies that the manifest was signed by info@projectious.work through GitHub’s OIDC issuer. Require that identity check in controlled environments with:

sh
curl -fsSL https://raw.githubusercontent.com/projectious-work/ainfra/v1.x-release/scripts/install.sh \
  | VERIFY_SIGNATURE=1 bash

Alternatively, download the archive, checksum manifest, and signature bundle from the release page and verify them before placing ainfra on PATH.

To build from source instead, install Go 1.26.5, clone the repository, and run:

sh
go build -o ainfra ./cmd/ainfra
./ainfra version

Git is required when locking or updating a Git template source. OpenTofu is required for planning and infrastructure changes; Ansible is required for configuration and convergence operations. Run ainfra doctor environment after installation to see which tools the intended workflow needs.

The repository also contains a convenience Dockerfile; no prebuilt image is published. Build the Linux binaries and local image with:

sh
scripts/build-targets.sh dist
docker build --build-arg TARGETARCH=amd64 -t ainfra:local .
docker run --rm ainfra:local version

Configuration

ainfra combines immutable, typed configuration layers without changing a deployment or template. From lowest to highest precedence, the layers are:

  1. compiled defaults;
  2. system configuration;
  3. user configuration;
  4. deployment-local ainfra.config.yaml;
  5. an explicit --config file or AINFRA_CONFIG; and
  6. supported environment variables and command options.

On Linux, the system and user files are /etc/ainfra/config.yaml and $XDG_CONFIG_HOME/ainfra/config.yaml (falling back to ~/.config/ainfra/config.yaml). On macOS they are /Library/Application Support/ainfra/config.yaml and ~/Library/Application Support/ainfra/config.yaml.

Files are strict YAML documents using apiVersion: ainfra.projectious.work/v1 and kind: Configuration. Unknown fields, unsupported versions, invalid enum values, and malformed environment values are rejected. Optional absent files are reported as absent; an explicitly selected missing file is an error.

Project discovery uses, in order, a positional target, --project, AINFRA_PROJECT, or the nearest ancestor containing ainfra.yaml. An explicit directory never searches descendants. A directory and its contained ainfra.yaml identify the same deployment.

Repository-controlled project configuration may set only UI choices and logging.level. It cannot select executables, storage paths, or log destinations. Native OpenTofu and Ansible variable files stay opaque and are never translated into ainfra configuration.

Use the environment doctor to inspect the effective values and provenance:

sh
ainfra doctor environment --format json
ainfra doctor environment --project path/to/deployment

The report shows each winning layer and overridden layers without displaying secret values.

ainfra template lock and ainfra template update use the same precedence for trusted acquisition controls. An absolute paths.cache selects private template storage, and executables.git selects the Git executable. Project configuration cannot control either value; such attempts fail closed. Git is resolved lazily, so local-only locking does not require it.

Operational logging#

Operational logs are separate from command output and retained run evidence. The default level is warn and the default destination is human-readable stderr. Increase verbosity with -v, -vv, or -vvv, or select an explicit --log-level error|warn|info|debug|trace; the two forms are mutually exclusive.

--log-format text|json selects the sink format independently of --format. --log-file ABSOLUTE_PATH adds a private rotating file sink, and --syslog adds the local system logger. File rotation defaults to 10 MiB, five retained files, seven days, and compression. Configuration files expose the detailed rotation and syslog settings.

All sinks receive the same structured event after exact-value and common credential-shape redaction. A requested sink that cannot be initialized or written is a command failure; ainfra does not silently discard an audit destination. ainfra logs never reads these operational sinks—it reads only the selected run’s retained lifecycle evidence.

Output, inventory, and Ansible

Validate OpenTofu output and configure hosts from a reviewed run.

This Phase 5 workflow is published in v1.0.0-alpha.5.

Phase 5 keeps the OpenTofu-to-Ansible handoff narrow. A template declares one standardized, non-sensitive OpenTofu output. ainfra validates only that value, converts it to deterministic inventory, and passes template-specific settings through native Ansible variable files rather than an ainfra variable model.

Run the stages separately#

Start with a successfully applied reviewed run:

sh
ainfra output ./deployments/example --run RUN_ID
ainfra inventory ./deployments/example --run RUN_ID
ainfra configure ./deployments/example --run RUN_ID
ainfra configure ./deployments/example --run RUN_ID --check

output writes private output.json; inventory derives private inventory.yaml from that validated artifact. Normal configure performs the declared playbook. The separate --check execution verifies that every expected host was processed with zero changes, failures, and unreachable hosts.

Use --format json for the versioned machine result. Results contain only artifact digests, engine attribution, and evidence paths. Ansible Runner artifacts remain marked sensitive.

Run the composed pipeline#

deploy applies the exact reviewed plan and then runs every applicable Phase 5 stage in order:

sh
ainfra deploy ./deployments/example --plan RUN_ID

It never creates an implicit plan. Its JSON result includes ordered stage reports for output, inventory, configure, and configure-check. Infrastructure- only templates record all four stages as not-applicable, skip Ansible, and do not require ansible-runner to be installed.

SSH trust#

SSH inventory requires spec.ssh.knownHosts to name a populated native file that was bound into the reviewed run. ainfra forces host-key checking and passes the bound snapshot to Ansible. It does not discover, accept, or repair host keys automatically.

Recovery rules#

Output validation and inventory conversion are non-mutating and may be rerun. A normal or check-mode configuration stage starts at most once for a retained run. If it fails or is interrupted, inspect its Runner artifacts and run evidence before creating a new reviewed plan; ainfra does not automatically repeat a potentially mutating playbook.

Hetzner private K3s template

Phase 9 ships the live-certified provider-backed production candidate at templates/hetzner-kubernetes-baseline in v1.0.0-alpha.9.

It combines ownership-labelled Hetzner infrastructure, private management addresses, initial pinned K3s bootstrap, an externally managed Cloudflare Tunnel, and an optional temporary SSH bastion. It does not install workloads or provide day-two Kubernetes operations.

Security boundary#

  • Node firewalls expose no public inbound TCP service.
  • Generated inventory contains private addresses only.
  • The bastion defaults off and requires explicit narrow CIDRs.
  • Host keys must be verified out of band and recorded in the deployment’s bound known_hosts file; host-key checking stays enabled.
  • Initial private-node configuration uses an operator-owned SSH ProxyJump route through the temporary bastion. Remove the bastion only after verifying the tunnel and an independent private management path.
  • SSH private keys, K3s tokens, and tunnel tokens remain external.
  • K3s and cloudflared binaries require architecture-specific SHA-256 values.

Evaluation#

The Phase 9 certification exercised a disposable three-node control plane, verified K3s health and idempotent convergence, confirmed Cloudflare Service Auth SSH after temporary-bastion removal, executed the exact reviewed destroy plan, and independently verified that no owned Hetzner resources remained. That evidence validates the released template candidate; it does not authorize or certify a future deployment automatically.

Run ainfra doctor template templates/hetzner-kubernetes-baseline and the template’s tests/validate.sh for offline conformance. The test performs OpenTofu initialization/validation, Ansible syntax, security policy, and native-variable documentation drift checks.

Live use additionally requires explicit cost approval, protected provider credentials, a reviewed apply/configure/check cycle, reviewed bastion removal, reviewed destroy, and an ownership-scoped Hetzner API query proving that every owned resource is gone. Offline checks are not certification evidence.

Usage

The current alpha CLI validates local contracts, diagnoses deployments, resolves immutable local or Git template sources, creates and executes reviewed OpenTofu plans, configures hosts through Ansible, and retains sanitized lifecycle evidence. Static help and version inspection do not perform project discovery, network access, or child-tool execution:

sh
ainfra help
ainfra help version
ainfra version
ainfra --format json version

Validate and diagnose local state with:

sh
ainfra init example-deployment
ainfra doctor environment
ainfra doctor deployment example-deployment
ainfra doctor template path/to/materialized-template
ainfra doctor all example-deployment

Resolve and bind template content with:

sh
ainfra template lock example-deployment
ainfra template update example-deployment

Both mutations accept --config, follow trusted configuration precedence, and emit stable text or JSON results. template lock creates the initial binding and refuses to replace a changed lock. template update is the explicit path for accepting a new source identity, Git commit, or tree digest. Doctor never resolves a mutable Git ref and remains read-only unless guarded local reconciliation is explicitly requested.

Create a private, immutable saved plan and apply that exact reviewed plan with:

sh
ainfra plan example-deployment
ainfra apply example-deployment --plan RUN_ID

The plan result supplies RUN_ID and the complete follow-up command. Apply refuses missing, destroy-intent, replayed, or stale plans and never creates an implicit replacement plan. See Reviewed plans for binding, evidence, JSON-output, and interruption details.

Destruction requires its own reviewed plan and exact run ID:

sh
ainfra plan example-deployment --destroy
ainfra destroy example-deployment --plan RUN_ID

Inspect retained evidence and ambiguous recovery state without reading OpenTofu state:

sh
ainfra status example-deployment
ainfra logs example-deployment --run RUN_ID
ainfra doctor run example-deployment

The published v1.0.0-alpha.7 includes reviewed destroy, recovery, retained logs, and operational logging in addition to output, inventory, Ansible, and composed deploy commands; see Output, inventory, and Ansible. Phase 7 adds guarded MCP serving. The interface is read-only by default and fixes one project root for the lifetime of the process:

sh
ainfra mcp serve --stdio --project /path/to/deployment

Saved plan creation is discoverable only when explicitly enabled:

sh
ainfra mcp serve --stdio --project /path/to/deployment \
  --capability planning

Enabling a capability does not approve a lifecycle mutation. Deployment and destruction require independent, operation-bound authorization.

Deployment capability startup requires AINFRA_MCP_APPROVAL_KEY containing at least 32 bytes of verifier key material. Each mutation request must carry an externally issued ainfra.approval/v1 artifact signed with HMAC-SHA256. The artifact binds the canonical root, operation, saved plan ID and digest, intent, caller, independent approver, issue time, expiry, and nonce. ainfra exposes no approval-signing command or MCP tool.

The default registry exposes read-only diagnostics, status, sanitized retained output and inventory, all published v1 schemas, and the normative MCP server contract. It never exposes raw engine streams.

With an authorization provider configured, deployment exposes apply, configure, convergence-check, and deploy operations. destruction additionally exposes exact reviewed destroy-plan execution and cannot be enabled without deployment. Every stdio request frame is limited to 1 MiB.

Tool and resource execution is bounded to eight concurrent requests. Client cancellation propagates into planning and lifecycle operations, including any permitted child process through the existing application execution contract.

Guarded MCP server

ainfra can serve its typed application operations to MCP clients over stdio. The server is bound to one deployment at startup, is read-only by default, and never treats tool annotations or conversational claims as authorization.

Start a read-only server#

Run the server from a deployment directory:

sh
ainfra mcp serve --stdio

Or bind it explicitly:

sh
ainfra mcp serve --stdio --project path/to/deployment

Protocol frames are the only data written to stdout. Diagnostics and operational logs use stderr or the configured logging sinks. The startup project, configuration, executable paths, cache, run storage, environment allowlist, and logging destinations cannot be overridden by tool arguments.

The default registry provides read-only tools for:

  • build and contract versions;
  • project, deployment, and verified template inspection;
  • retained status, standardized output, and inventory;
  • deployment, run, template, and environment diagnostics; and
  • published schemas and contract resources.

Enable optional capabilities#

Additional tools must be allowlisted when the server starts:

sh
ainfra mcp serve --stdio --capability planning
ainfra mcp serve --stdio --capability deployment
ainfra mcp serve --stdio \
  --capability deployment \
  --capability destruction

planning adds reconciliation, template lock/update/migration, and reviewed OpenTofu planning. Planning may create private run evidence or digest-addressed template cache entries, but it does not apply infrastructure or publish a template lock.

deployment adds independently authorized reconciliation, template writes, apply, configure, and composed deploy. destruction adds destroy execution, requires deployment, and accepts only an approval with destroy intent. Enabling a capability makes tools discoverable; it does not approve a mutation.

Configure signed approvals#

Mutation tools require a trust file selected by the operator at startup:

sh
ainfra mcp serve --stdio \
  --project path/to/deployment \
  --capability deployment \
  --authorization-trust path/to/mcp-trust.json

The trust file must be a private, non-symlink JSON file:

json
{
  "schemaVersion": 1,
  "issuers": [
    {
      "id": "operator-1",
      "publicKey": "BASE64_ED25519_PUBLIC_KEY"
    }
  ]
}

An approval is a strict JSON envelope containing schemaVersion: 1, the raw grant object, and a base64 Ed25519 signature. Sign the UTF-8 bytes ainfra-mcp-authorization-v1\n followed by the exact raw JSON bytes used as the envelope’s grant value.

The grant binds all authority-relevant inputs:

json
{
  "authorizationId": "approval-2026-08-16-1",
  "issuer": "operator-1",
  "caller": "agent-1",
  "projectRoot": "/absolute/canonical/deployment",
  "operation": "apply",
  "planId": "REVIEWED_PLAN_ID",
  "planDigest": "sha256:...",
  "intent": "apply",
  "approvedAt": "2026-08-16T12:00:00Z",
  "expiresAt": "2026-08-16T12:05:00Z"
}

The issuer must differ from the caller. Approval must be current and match the startup root, canonical operation, plan identity, digest, and intent exactly. Opaque approval bytes are never logged, returned, or retained. Sanitized authorization evidence records only the authorization ID, caller, issuer, binding, and expiry.

Reviewed mutation workflow#

For infrastructure mutations, use the planning capability to create a saved plan and present its run ID and digest for independent review. Restarting with both planning and deployment is allowed, but the same agent request cannot approve the plan it created. Apply, deploy, and destroy always consume an existing reviewed plan; no mutation tool creates an implicit replacement.

Template lock/update and reconciliation previews similarly return stable plan IDs and digests. Their execution tools recompute the complete plan under the deployment operation lock before writing. Changed source content or stale preconditions require a new preview and approval.

Operational safety#

  • Keep trust files and private keys outside the deployment and source tree.
  • Enable only the capabilities needed for the current session.
  • Use short approval expiries and unique authorization IDs.
  • Treat an interrupted mutation as inspection-required; consult ainfra.status and retained logs before retrying.
  • Do not expose the stdio server through a network bridge. Network transports are outside the v1 threat model.
  • Close the MCP session normally so the stdio child can shut down cleanly.

The MCP result envelope is versioned independently from the transport. Clients should inspect apiVersion, tool, ok, result, and diagnostics rather than parse display text.

Templates

Templates are self-contained, reviewable infrastructure environments for specific AI-agent workloads. Phase 8 defines the released authoring contract; Phase 9 ships the first provider-backed production candidate after offline conformance and a cost-approved disposable live certification.

Hetzner private K3s production candidate#

templates/hetzner-kubernetes-baseline provisions private-management Hetzner nodes, bootstraps pinned K3s, and runs a connector for an externally managed Cloudflare Tunnel. Public SSH exists only through an opt-in temporary bastion with a narrow source allowlist.

The candidate has complete native-variable documentation, offline conformance coverage, and Phase 9 certification evidence. The certified lifecycle created three private K3s control-plane nodes, verified convergence and tunnel-only administration, removed the temporary bastion, destroyed the exact reviewed resources, and independently confirmed that no owned Hetzner resources remained. Every new deployment still requires its own authorization, trust bootstrap, validation, and teardown evidence.

Start by copying the reference template. Keep the applicable authoring layout intact:

text
README.md
docs/variables.md
tofu/{versions,variables}.tf
tofu/outputs.tf                           # when the template emits outputs
ansible/                                  # when hosts are configurable
tests/{README.md,validate.sh,fixtures/}   # output fixture when inventory exists

The template manifest declares identity, engine requirements, inventory and input/output contracts. OpenTofu owns infrastructure resources; Ansible owns host configuration. Keep native tool configuration inside those engines rather than duplicating it in ainfra.

Authoring workflow#

  1. Copy the reference template and give the manifest a unique name and version.

  2. Declare OpenTofu variables, outputs and required provider versions.

  3. Add Ansible inventory/playbooks and map host facts deliberately.

  4. Document every supported input in docs/variables.md.

  5. Write the lifecycle guide in README.md, including prerequisites, architecture, network exposure, costs, failure handling and teardown.

  6. Create a clean-room fixture and make tests/validate.sh exercise it.

  7. Run the local conformance gate and native validation:

    sh
    ainfra doctor template . --format json
    ./tests/validate.sh

ainfra doctor template validates the portable authoring contract. It reports native OpenTofu and Ansible validation as delegated, so the clean-room script must run compatible tools itself. A passing doctor report is not authorization to create billable resources; obtain explicit approval immediately before any live provider lifecycle.

Documentation contract#

docs/variables.md contains these H2 sections, in order: OpenTofu variables, Ansible variables, Cross-variable rules, and Examples. The variable tables identify name, type or shape, required/default values, constraints, sensitivity and description.

The template README.md covers prerequisites, architecture, variables, network, cost, failure, teardown, compatibility and validation. Never put a credential in the manifest, fixture, README or generated evidence.

For a complete machine-oriented checklist, see the AI template-authoring guide. The normative schema and lifecycle details remain in the template contract.

AI template-authoring guide

This is the self-contained execution entry point for an AI agent authoring an ainfra v1 template. It is a conformance workflow, not permission to provision a live environment.

Objective#

Produce a portable template whose ainfra documents, native OpenTofu and optional Ansible content can be validated from a clean room. Do not invent an ainfra variable language, generate credentials, contact a provider during local validation, or claim live support without disposable lifecycle evidence.

Required deliverables#

Create this minimum tree. Omit tofu/outputs.tf, ansible/, ansible-vars.yaml and the output fixture when the manifest declares inventory: none and no other output is needed:

text
ainfra-template.yaml
README.md
docs/variables.md
tofu/versions.tf
tofu/variables.tf
tofu/outputs.tf
ansible/
examples/minimal/ainfra.yaml
examples/minimal/terraform.tfvars
examples/minimal/ansible-vars.yaml
tests/README.md
tests/validate.sh
tests/fixtures/output.json

The complete provider-free reference template is an example, not hidden context required to use this guide.

Manifest#

Use this infrastructure-only starting point and change the name, version and engine constraint deliberately. The containing directory must have the same name as metadata.name.

yaml
apiVersion: ainfra.projectious.work/v1
kind: Template
metadata:
  name: example-template
  version: 1.0.0
spec:
  engines:
    tofu:
      directory: tofu
      version: ">=1.10.0 <2.0.0"
  outputs:
    inventory: none

For configurable hosts, add an Ansible engine with directory, playbook and an explicit version range, then set inventory to the exact OpenTofu output name. The accepted manifest fields are defined by the published schema; unknown fields are invalid.

Native variables and dependencies#

Declare infrastructure inputs in tofu/variables.tf and host-configuration inputs in native Ansible variable files. ainfra passes deployment-selected native files through unchanged. Pin OpenTofu providers and modules; commit .terraform.lock.hcl whenever provider selections exist. Pin every external Ansible collection and role in requirements.yml.

docs/variables.md contains these H2 sections in this exact order:

  1. OpenTofu variables
  2. Ansible variables
  3. Cross-variable rules
  4. Examples

Each engine section uses:

NameType or shapeRequiredDefaultValid values and constraintsSensitiveDescription
example_namestringyes—Non-emptynoOperational effect.

Document every deployment-facing native variable. Reproduce types, defaults, validation and sensitivity accurately, and explain replacement, access, cost or teardown consequences. Examples must be minimal, valid and non-secret.

Standard output#

An inventory-producing template emits one non-sensitive OpenTofu output whose name matches the manifest. Its JSON shape is:

json
{
  "schema_version": "1",
  "hosts": {
    "node": {
      "groups": ["all"],
      "connection": {"type": "local"}
    }
  }
}

Use only the connection types and fields allowed by the standard-output schema. Outputs must contain stable routing facts, never credentials or private keys.

README and security#

The template README documents purpose and unsupported outcomes, provider and account prerequisites, least-privilege credentials, an architecture diagram, variables, network/access exposure, costs and billable opt-ins, backend/state ownership, the full lifecycle, equivalent native commands, failure/recovery, teardown with independent verification, compatibility and dated validation limitations.

Default public ingress off. Keep management addresses private where supported, make any temporary administrative ingress explicit and narrow, verify SSH host keys independently, label owned resources, keep credentials external, and make destroy cover every resource the template owns. Fixtures and evidence contain no secret value.

Local validation#

tests/validate.sh is executable and runs, as applicable:

sh
tofu -chdir=tofu fmt -check
tofu -chdir=tofu init -input=false
tofu -chdir=tofu validate
ainfra doctor template . --format json
./tests/validate.sh

The script also checks dependency pins, documented-variable drift, secret policy, standard output fixtures, Ansible syntax, and a zero-change check-mode pass. The doctor reports native checks as delegated; its skip is an instruction to run the script, not a conformance success.

Disposable live acceptance#

Stop before this section until a human gives explicit, immediate approval for the named provider account, region, cost ceiling and teardown window. After approval: create and review a plan; apply exactly that plan; configure; verify zero-change convergence; reapply only if the template requires that proof; create and review a destroy plan; destroy; and independently query the provider to prove owned resources are absent.

Retain sanitized evidence containing template revision/digest, date, operator, provider and tool versions, approved scope/cost ceiling, plan summaries, configure/check results, destroy result, independent teardown confirmation, limitations and final pass/fail. Never retain state, raw credentials or secret values in published evidence.

Completion checklist#

  • Manifest identity, version, supported tools and input/output contracts are accurate.
  • Provider/module and Ansible dependencies are pinned and applicable native lockfiles are committed.
  • The README explains prerequisites, architecture, variables, network, cost, failure behavior, teardown, compatibility and validation.
  • docs/variables.md has the required headings and complete tables.
  • A clean-room fixture exists and tests/validate.sh validates it.
  • Variable drift, secret policy and standard output fixtures are tested.
  • ainfra doctor template . --format json has no fail findings.
  • ./tests/validate.sh passes using compatible native tools.
  • Live support is claimed only after the approved disposable lifecycle, independent teardown verification and sanitized evidence are complete.

Reviewed OpenTofu Plans

ainfra separates OpenTofu planning from infrastructure mutation. plan creates a private saved plan and immutable review record. apply requires the exact run ID and executes only those saved bytes.

Prerequisites#

The deployment must have a valid ainfra.yaml, a current ainfra.lock, and a verified template cache entry. OpenTofu is discovered only when plan or apply runs. Pin an explicit executable in trusted user or explicit configuration when required:

yaml
apiVersion: ainfra.projectious.work/v1
kind: CLIConfig
executables:
  tofu: /usr/local/bin/tofu

Repository-controlled configuration cannot select executable, cache, or run paths.

Create and review a plan#

sh
ainfra plan path/to/deployment

Planning creates a collision-resistant run ID and an owner-only directory under the configured runs root. It snapshots declared backend and variable files, materializes the locked template, runs tofu init, creates plan.tfplan, and reduces tofu show -json to structural action counts. Raw plan values are not persisted in the summary.

For automation, request the stable v1 result:

sh
ainfra plan path/to/deployment --format json

The result identifies apply versus destroy intent, deployment and template digests, the aggregate input digest, saved-plan digest, structural engine report, evidence paths, and exact next command.

Destroy planning is explicit:

sh
ainfra plan path/to/deployment --destroy

It produces a destroy-intent record and a ainfra destroy ... --plan RUN_ID next command. ainfra apply always refuses that record. After reviewing the structural delete counts, execute only that saved plan:

sh
ainfra destroy path/to/deployment --plan RUN_ID

ainfra repeats every binding check described below and invokes tofu apply <saved-destroy-plan> as an argument array. It never runs tofu destroy, creates an implicit replacement plan, or reads OpenTofu state.

Apply the exact reviewed plan#

sh
ainfra apply path/to/deployment --plan RUN_ID

ainfra acquires the deployment operation lock and then reverifies every plan binding: deployment manifest, ordered native inputs, template lock and cache, template-controlled workspace bytes, OpenTofu executable and version, summary, and saved plan. Any drift returns the stale-binding exit code without starting OpenTofu.

After recording durable started evidence, ainfra invokes only:

text
tofu apply -input=false -no-color <contained-relative-plan.tfplan>

There is no implicit replanning. A second invocation with the same run ID is refused, including when another apply process is already using it.

Evidence and recovery#

events.jsonl records the started event before execution and a succeeded, failed, or cancelled terminal event afterward. Cancellation or an ambiguous process interruption additionally records inspection-required and disables automatic retry.

Destroy uses the same evidence and interruption boundary. A zero OpenTofu exit code records the engine result; it is not independent proof that the provider contains no owned resources. Follow the certified template’s provider-side teardown procedure after every destroy.

A certified provider template’s teardown procedure must identify:

  • the authenticated, read-only provider API or inventory command;
  • the deployment ownership labels, account, project, and region boundaries;
  • the paginated query and the exact empty-result condition;
  • a bounded wait for provider eventual consistency;
  • the sanitized evidence retained for certification; and
  • an emergency cleanup and escalation path when resources remain.

The verification must query the provider directly. It must not infer absence from the destroy exit code, a saved plan, ainfra events, or direct inspection of OpenTofu state.

When inspection is required, do not create or apply another plan merely to clear the error. Preserve the run directory, inspect it with:

sh
ainfra doctor run path/to/deployment

Then inspect the backend and infrastructure using the provider’s native, read-only facilities before deciding whether a new plan is safe.

Saved plans are sensitive even though the structural summary is sanitized. Do not copy run directories into source control, attach them to public issues, or parse plan.tfplan outside the trusted local environment.

Retained status and logs#

ainfra status DEPLOYMENT derives lifecycle state only from private run records and events. It does not inspect OpenTofu state. Interrupted mutations are inspection-required and are never marked safe for automatic retry.

ainfra logs DEPLOYMENT --run RUN_ID shows the sanitized ainfra event timeline. Add --errors to select typed failed, cancelled, and inspection-required states; it never searches localized prose. Raw child evidence requires an explicitly retained stream, child source, stream, and interactive confirmation:

sh
ainfra logs DEPLOYMENT --run RUN_ID --source opentofu \
  --raw --stream stderr

Automation must add both --non-interactive and --yes. Raw access is incompatible with --errors and JSON output. Selected bytes go directly to stdout and the warning goes to stderr; raw bytes never enter the result envelope or normal renderer. Stream files must be private regular files inside the selected run.

For Ansible-backed runs, --source ansible-runner reads the retained native Runner v2 job-event artifacts. --errors selects only native failure, unreachable, asynchronous-failure, and error event types. Words such as ERROR in localized stdout never determine classification, and event data is not copied into the sanitized display. Unsupported Runner protocol majors, symlinks, public files, excessive trees, and oversized artifacts fail closed.

References

Doctor commands#

text
ainfra doctor [TARGET]
ainfra doctor all [TARGET]
ainfra doctor environment
ainfra doctor deployment [TARGET]
ainfra doctor template [LOCAL_TEMPLATE]
ainfra doctor run [TARGET]

Bare doctor is an exact alias of doctor all. The aggregate selects the applicable environment, deployment, already-resolved template, and latest-run checks. Missing optional evidence is skip, never pass.

Every finding has a stable code, scope, check ID, status, severity, message, component or path, and reconciliation state. Failed and skipped checks also provide a next action. --format json emits exactly one ainfra.result/v1 object on stdout; diagnostics and confirmation prompts do not contaminate JSON stdout.

Local reconciliation#

--reconcile can repair only registered, ainfra-owned local artifacts. The currently registered repair creates the deployment .ainfra directory or restricts it to owner-only mode (0700). It does not edit manifests, native inputs, templates, state, credentials, or remote infrastructure.

Interactive use prints the complete plan and asks for confirmation. Automation must provide all three options:

sh
ainfra doctor deployment path/to/deployment \
  --reconcile --non-interactive --yes

ainfra locks the deployment, rechecks plan preconditions, applies eligible actions, and reruns affected checks. Failed actions remain in the result as failed or still_failing; stale plans are refused before writing.

Exit codes#

CodeMeaning
0Successful command; warnings and skips may still be present.
1Operation, child-tool, rendering, or reconciliation failure.
2Invalid invocation, configuration, or input contract.
3Missing or incompatible required dependency.
4Security or trust refusal.
5Stale reviewed-plan binding (reserved for lifecycle commands).
6Interrupted or ambiguous mutation (reserved for lifecycle commands).

Machine-output schemas and examples are maintained under spec/schemas/v1/ and spec/examples/v1/machine-output/.

Tutorials

Validate a template in a clean room#

Copy the reference template, make only the required local edits, then run:

sh
ainfra doctor template . --format json
./tests/validate.sh

The first command validates the portable authoring contract. The test script runs the compatible native tool checks and fixture assertions. Do not create cloud resources as part of this tutorial: live lifecycle work requires explicit approval at the point it is performed.

See the template authoring guide for the required layout and the AI guide for an execution checklist.

Contributing

Contributions to ainfra should begin with the repository’s contribution and security guidance.

The complete development workflow is maintained in the repository CONTRIBUTING.md.

The release command boundary is strict. All release lanes must first point to one exact commit; release-freeze records that commit and tree before any artifact production:

Do not paste the complete sequence into one terminal. HOST commands run in the normal host checkout; DEVCONTAINER commands run in that checkout’s ainfra devcontainer. Both environments see the same repository files. No second clone, checkout, worktree, or artifact copy is needed, and Go is never installed or invoked on the host.

  1. HOST: run release-freeze after all merges and lane promotions.
  2. DEVCONTAINER: run release-package. This is the only step that uses Go. It cross-builds four archives and uses Syft to create their SBOMs. A missing go error means the command was run on the host; return to the devcontainer instead of installing Go on the host.
  3. HOST: run release-host-prepare. It needs Git and Python, verifies the packaged artifacts, and does not compile anything.
  4. HOST: run release-host. It uses uv, Docker, Syft, and Grype to test and inventory the independently built container image.
  5. HOST: run release-publish --dry-run to preflight credentials and publication.
  6. HOST: run release-sign, followed by resumable release-publish.

Syft therefore belongs in both environments, but it inventories different subjects. The workspace pins an aibox release whose catalog supplies the complete supply-chain toolset: Gitleaks, OSV-Scanner, Syft, Grype, and Cosign. Run aibox apply and rebuild the devcontainer after changing that pin or its tool selections. Agents must never receive the host Docker socket or equivalent container-runtime authority. See the repository guide for the complete commands, prerequisites, and checks.

Roadmap

The ainfra v1 roadmap begins with the core lifecycle and expands through hardening, templates, integrations, and release readiness. It is generated directly from the authoritative specification.

Total
23
Shipped
10
Idea
13

Confidential infrastructure

attested platforms and confidential secret release

Phase 22 ◇ Idea

Confidential computing and live attestation

Provision and describe confidential-computing targets, supported TEE and runtime capabilities, Trustee and KBS services, attestation and reference-value policy, confidential secret backends, and signed live or remote evidence; prove an ainfra-to-aibox handover that binds deployment provenance to current attested infrastructure without making ainfra the workload or secret-management authority.

Engine evolution

deliberately distant architectural options

Phase 21 ◇ Idea

Exchangeable infrastructure engine

Evaluate a versioned provisioning-engine contract that could support alternatives to OpenTofu, including configuration-led or Ansible-only templates, without weakening native inputs, reviewed change approval, teardown semantics, recovery, or audit evidence.

Policy and coordination

potential directions

Phase 20 ◇ Idea

Remote execution protocol

Run the same CLI lifecycle in a controlled remote environment while preserving explicit credentials, reviewed plans, child-tool boundaries, and auditable results.

Phase 19 ◇ Idea

Deployment-set coordination

Coordinate several independent deployments while preserving separate state, locks, reviewed plans, approvals, and failure boundaries.

Phase 18 ◇ Idea

External policy checks

Pass sanitized manifests or plan summaries to an external policy engine at explicit lifecycle gates and enforce its result without owning a policy language.

Extended provisioning

infrastructure-adjacent bootstrap

Phase 17 ◇ Idea

External cluster bootstrap

Invoke a dedicated external tool for initial k3s or comparable Kubernetes installation through a bounded child-process contract, like OpenTofu and Ansible, without taking on cluster maintenance.

Template ecosystem

authoring, compatibility, and trust

Phase 16 ◇ Idea

Authoring and editor integration

Publish schema bundles, completions, and editor integrations that consume doctor JSON without introducing another configuration language.

Phase 15 ◇ Idea

Operational evidence export

Export sanitized and attributable deployment evidence for audits and support without exporting plans, state, credentials, or sensitive output.

Phase 14 ◇ Idea

Catalog and trust metadata

Define interoperable discovery metadata for certified Git-hosted templates, including compatibility, publisher identity, signatures, attestations, and revocation.

Phase 13 ◇ Idea

Compatibility and migration experience

Improve fixtures, compatibility diagnosis, deprecation reporting, and safe template-contract migration previews as contracts evolve.

Template coverage

useful deployments out of the box

Phase 12 ◇ Idea

Specialized compute templates

Add templates for Vast.ai and similar GPU or on-demand compute providers, with explicit lifecycle and teardown limitations.

Phase 11 ◇ Idea

Complete major-cloud templates

Add comparable certified templates for AWS, Microsoft Azure, and Google Cloud while preserving the same native-file and output contracts.

Phase 10 ◇ Idea

Verifiable consumer target handover

Evolve the single standardized infrastructure result with a versioned, non-secret target projection for workload consumers; cover capabilities, endpoints, symbolic secret-provider, credential, and trust references, compatibility, deterministic deployment provenance, DSSE signing, signer and freshness policy, and one proven ainfra-to-aibox handover without treating discovery or signatures as proof of current live infrastructure state.

Phase 9 ✓ Shipped

Initial production template

Ship and certify the first Kubernetes-ready template with private access, tunneled ingress, and temporary administrative access.

Read phase note →

Template authoring

the v1 authoring promise

Phase 8 ✓ Shipped

Template authoring and conformance

Publish schemas, examples, conformance diagnostics, and human and AI guidance, then prove through clean-room authoring that a new template can be created without reading ainfra implementation source.

Read phase note →

Automation interface

read-only by default with explicitly authorized agent operations

Phase 7 ✓ Shipped

Guarded MCP server mode

Serve doctor, inspection, status, schemas, and sanitized run results over MCP stdio by default, and expose explicitly allowlisted planning and plan-bound lifecycle mutations through the same typed application use cases, authorization, locking, evidence, and recovery contracts.

Read phase note →

Infrastructure lifecycle

opentofu and ansible orchestration

Phase 6 ✓ Shipped

Destruction, recovery, and hardening

Add reviewed destroy plans, teardown verification, interruption recovery, redaction, binding checks, cache defenses, and negative security fixtures.

Read phase note →
Phase 5 ✓ Shipped

Output, inventory, and Ansible

Validate standardized non-secret output, generate deterministic inventory, run Ansible with native variables, and verify zero-change convergence.

Read phase note →
Phase 4 ✓ Shipped

Reviewed OpenTofu plans

Initialize native OpenTofu configuration, create bound saved plans, apply only the reviewed plan, and record durable run evidence.

Read phase note →

Foundation

the product contract

Phase 3 ✓ Shipped

Immutable template sources

Resolve local and Git-subdirectory sources, lock revisions and digests, materialize contained workspaces, and detect source or cache drift.

Read phase note →
Phase 2 ✓ Shipped

Contracts and doctor

Discover deployments, parse manifests and native files, validate ainfra-owned contracts, and report stable text and JSON diagnostics plus safe local reconciliation through one doctor surface.

Read phase note →
Phase 1 ✓ Shipped

Go project foundation

Establish the Go module, command shell, typed results, process and filesystem boundaries, test fixtures, developer tools, and Linux/macOS builds.

Read phase note →
Phase 0 ✓ Shipped

Product specification

Finalize the v1 boundary, native-file contracts, security model, Go architecture, schemas, examples, and acceptance journeys.

Read phase note →