15 KiB
name, description
| name | description |
|---|---|
| apsuflow | Use when operating, deploying, configuring, troubleshooting, or explaining Apsuflow. Build a supported local, multi-environment, or cluster workflow from current requirements and CLI help; explain .apsu, networking, quorum, and backup/restore; refuse direct SQLite, runtime, or host-network mutation. Do not use for implementing Apsuflow, changing requirements, or generic container architecture advice. |
Operate Apsuflow
One job
Turn an operator goal into the supported path:
select target -> validate desired state -> inspect diff -> apply -> verify -> recover through public controls
This skill prevents agents from bypassing the control plane by editing SQLite, invoking the OCI runtime, or changing derived networking.
Authority
Before giving commands or causing effects, check only what affects the answer:
REQUIREMENTS.md— current product obligations;apsuflow <command> --help— installed command contract;docs/getting-started.md,docs/cli.md,docs/solo-vs-cluster.md,docs/operations.md, anddocs/grammar.md— explanations;examples/— syntax examples, never higher authority.
Follow requirements and observed CLI behavior when they conflict with prose, and
report the stale surface. SKILL.md is derived guidance, not authority. Update it
in the same change when applicable requirements or CLI behavior change.
Mental model
- One server owns desired state, reconciliation, authentication, and the API.
- Solo mode is one loopback-bound server without cluster mesh overhead.
- Cluster servers replicate control-plane state; agents execute workloads but do not add quorum votes.
- One signed
apsuflowexecutable runs on every supported Linux profile; profiles contain only host configuration and service scripts. The daemon never detects a distribution, libc, or init system. The control-plane server runs as the lockedapsuflowuser with zero effective Linux capabilities and access only to its state directory and listener. A separately supervised, capability-bounded root agent owns OCI, namespace, cgroup, CNI, firewall, host-path, and host-network effects. Missing agent prerequisites leave workloads stopped; the server never becomes an execution fallback. Both release-supported profiles prove a real agent-owned OCI workload, independent server and agent restart, and encrypted control-plane backup/restore. .apsucontains desired workloads and infrastructure, never secret values or mutable runtime state.- Services are continuously reconciled. Jobs are finite run-to-completion work. Application queues, message retries, DLQs, and workflow state belong in broker-backed workloads.
Infrastructure as code is central
Treat reviewed, version-controlled .apsu plus the selected context as the
source of production intent. Put workloads, provider pools, placement, routes,
health checks, resources, update policy, and sealed-name references there; keep
secret values and credentials outside Git. A production change starts as a
declaration diff, passes fmt --check, validate, and preferably apply --dry-run, then applies that exact reviewed file set. Record its revision,
context, dry-run, and post-apply workload/route evidence.
Keep one canonical ordered declaration set in the elected IaC repository; it may
contain multiple .apsu files. A host-side .apsu is only an optional
convenience mirror: keep it at one predictable path, prove it byte-identical to
the reviewed revision before use, and remove temporary inspection or deployment
copies. Do not create a second source of truth through hand-edited runtime state
or an uncommitted mystery file.
If the elected declaration is missing, stop and recover its reviewed repository
revision or ask the owner for it. A control-plane backup may restore the server's
canonical desired state, but it is not a supported .apsu export; never generate
replacement IaC from a backup, SQLite, or inferred live state. A manual restart
may be an emergency diagnostic action, but the durable repair belongs in IaC and
must converge after re-apply.
Install and update
For an existing legacy root daemon, use apsuflow host migrate rather than
manually replacing the service. Prepare canonical declarations and separate,
encrypted backup archive/key restore evidence; run the command without
--commit first and review its state classification. Unknown state, an
unidentified supervisor, an existing split profile, invalid backup evidence, or
invalid persisted image references must stop before effects. If no unused
bootstrap token remains, provide a retained administrator token file and create
the node-bound solo-agent join token through the public CLI before shutdown;
supply them with --admin-token and --agent-token. Pass
--allow-secret-env and --acknowledge-recovery-risk when the reviewed
migration declaration requires those same explicit apply acknowledgments.
Commit as root only after review; it quiesces persisted services and reapplies
the exact declarations through public controls. Retain its rollback
directory until the split services, exact declaration reapply, and
application-level data checks pass. Never recursively chown workload data or
repair migration through SQLite, OCI, CNI, firewall, or namespace mutation.
Use only the exact release tag and artifact named by docs/getting-started.md.
Official bootstrap and Cloud-Init entry points require that exact version,
verify SHA256SUMS with the repository-pinned SSH release identity, verify the
selected archive hash, replace atomically, and retain apsuflow.previous; never
use a mutable latest URL. Keep the previous executable and state backup until the new one
passes version, doctor, workload diagnosis, and each operator-facing route.
For a private image registry, pipe the exact registry credential in its
username:password form into apsuflow registry login <host> --password-stdin;
the command has no separate username flag. Do not put the credential in .apsu,
process arguments, shell history, a runtime auth file, or host-side container-tool
configuration. Login stores sealed _registry/{host} state, and joined agents
receive only their per-node-resealed pull credential during assignment. Establish
and verify this login before the first apply, then confirm the digest-pinned pull
and workload health through Apsuflow. If a known-good credential still yields
unauthorized, stop: that is an Apsuflow credential-delivery defect, not permission
to use skopeo login, copy an OCI archive, invoke the runtime, or otherwise bypass
the control plane.
For host-network workloads, confirm the old container released its host ports and the successor is the sole running generation; repeated replacement failures are a rollback signal, not a reason to mutate runtime state directly. Every successor release must carry proved upgrade and rollback evidence against its preceding supported artifact. Apsuflow is not published to a Cargo registry.
Local operation
apsuflow fmt --check app.apsu
apsuflow validate app.apsu
apsuflow diff app.apsu
apsuflow apply app.apsu
apsuflow status
apsuflow diagnose web
Use fmt without --check only when reformatting is intended. Use apply --prune only when declarations omitted from all supplied files should be
removed. Validation reports portability limits before mutation. Capabilities
whose failure can require manual recovery—host networking, devices, and writable
host paths—also require apply --acknowledge-recovery-risk; read the reported
source location and consequence before accepting it. A volume source is itself
an exact hard placement requirement; do not duplicate it in .apsu constraints.
On each execution host, list the roots the agent may provide, one absolute path
per line, in root-owned mode-0600 /etc/apsuflow-agent/host-paths, then restart
the agent. A listed root grants authority; it need not pre-exist. The agent
creates a missing volume directory only below a listed root and rejects paths
outside those roots or through symlinks. Container entrypoints initialize
application content, not host placement. Verify the actual route or job result;
a successful apply alone proves only admission.
The server daemon runs unprivileged without runtime capabilities; the separately
supervised agent remains rootful only for its authorized runtime and host effects.
Do not run client commands with sudo: protect server state and use an operator
context or token. Join credentials are node-bound, role-bound, single-use, and
expire after 15 minutes; never reuse, commit, or broaden them. server --no-auth
is acceptable only for disposable loopback use.
Multiple environments
A context selects an endpoint and credentials; it is not a namespace. Use separate servers/clusters for environments requiring isolation.
apsuflow context set --name dev --server https://dev.example:7654
apsuflow context set --name prod --server https://prod.example:7654
apsuflow context list
apsuflow status --context dev --token "$APSUFLOW_DEV_TOKEN"
Prefer protected secret injection and explicit --context in automation; avoid
putting tokens in shell history or committed files. If storing a token in an
Apsuflow context with context set --token, ensure the config file is readable
only by its owner. Keep shared and environment-specific
inputs as ordinary files:
apsu/base.apsu
apsu/env/dev.apsu
apsu/env/prod.apsu
Use the same ordered set for validate, diff, and apply:
apsuflow validate --context dev apsu/base.apsu apsu/env/dev.apsu
apsuflow diff --context dev apsu/base.apsu apsu/env/dev.apsu
apsuflow apply --context dev apsu/base.apsu apsu/env/dev.apsu
A later duplicate declaration replaces the earlier declaration completely. Do not add templating, inline secrets, or database edits to simulate environments.
.apsu essentials
Top-level declarations are service, job, and infra:
service web {
image: "nginx:1.25"
replicas: 2
cpu: 0.25
memory: 128Mi
route { listen: 80 expose: 8080 protocol: tcp }
}
job backup {
image: "ghcr.io/example/backup@sha256:<digest>"
command: ["/app/backup"]
restart: on-failure
max_retries: 2
schedule { trigger: cron "0 3 * * *" concurrency: forbid }
}
- Use markerless canonical syntax and let
fmtnormalize it. - Prefer digest-pinned production images.
- Reference secret names and set values with
apsuflow secret set; never embed secret values in.apsu. - Routes are L4 TCP/UDP forwarding, not HTTP hostname routing or application TLS.
- Writable host paths are node-local; their volume sources imply placement, the
agent's root-owned
host-pathsfile declares availability, and data needs a separate backup. - Use a service plus an external broker for consumers; use a job for bounded migration, backup, maintenance, or scheduled work.
- Scheduled declarations create job runs; inspect them with
apsuflow job runs NAME. Useapsuflow runfor an intentional ad-hoc foreground container; it does not create durable scheduled-job history.
Cluster and network
Start the first server on a stable address. Create a 15-minute, single-use join
credential bound to the exact node ID with apsuflow token create --name NODE --role server, then start that server with apsuflow server --peer FIRST:7946 --token TOKEN. For an agent, mint the same node-bound credential with --role agent and run apsuflow agent --node-id NODE --join FIRST:7654 --token TOKEN.
The --node-id value must exactly match the credential name. Never reuse or
commit a join credential.
Allow between nodes only required control-plane traffic:
- TCP 7654 — API;
- UDP 7946 — membership;
- UDP 51820 — WireGuard overlay.
Expose workload route ports separately. Do not manually create WireGuard peers,
CNI networks, nftables rules, or DNS records. Verify with node list, status,
containers, routes, and diagnose. Before infrastructure apply, inspect
status --json: the configured provider must be compiled and the needed effect
must not appear in unavailable_capabilities; resolve any
leaked_infrastructure_resources before retrying destructive operations.
Quorum
Use three server-role nodes when control-plane failure tolerance matters. Writes need a majority: one server has no redundancy, two require both, and three can lose one. Agents do not count toward quorum.
On quorum loss, preserve committed state and reject unsafe writes. Use explicit
promotion only for its documented recovery case. store force-leader --confirm
is a last resort after the old writer is proven unavailable and possible loss is
accepted; scripts must acknowledge the committed writer term and sequence. It is
not a way to clear an error.
Backup and recovery
Create a client-owned encrypted control-plane backup and move both outputs off-cluster, storing the key separately from the archive:
apsuflow backup --context prod \
--key-out prod-control-plane.key \
prod-control-plane.tar.zst
Regularly prove restore against an empty, disposable target:
apsuflow restore --context recovery \
--key prod-control-plane.key \
prod-control-plane.tar.zst
apsuflow status --context recovery
This backs up Apsuflow control-plane state, not bind-mounted application data, databases, object stores, or brokers. Back those up natively and test both paths. Restore replaces authentication state and removes the empty target's bootstrap token because it cannot match restored credential hashes. Retain an original administrator credential outside the lost server data directory; without one, recovery is blocked rather than silently granting new access.
Never mutate store.db, WAL files, generations, assignments, certificates, or
sealed values with SQLite/SQL. Never invoke crun/youki or host-network tools to
force convergence. Read-only host inspection may support diagnosis, but effects
must return through Apsuflow or stop with a precise blocker.
Decision and refusal
- Identify the elected IaC revision, exact file order, context, and topology.
- Validate
.apsuinputs and secret references; run diff or dry-run before apply. - Ask before prune, restore, promotion, force-leader, or billable/destructive provider actions.
- Apply the exact reviewed declaration through CLI/API and prove convergence, workload behavior, and operator-facing routes.
- Diagnose drift through public status, logs, and recovery controls; repair the declaration rather than the runtime.
Do not invent built-in queues, DLQs, application TLS, HTTP hostname routing, or
distributed storage; use contexts as namespaces; put values/templates in .apsu;
use sudo for routine client calls when credentials suffice; or bypass desired
state through SQLite, runtime, CNI, WireGuard, nftables, or DNS mutation.
Report the target context/topology, .apsu inputs, effects, validation and
workload proof, backup impact, and any refused unsupported operation.