Operations runbook¶
Before market open¶
- Verify the deployed package hash and artifact attestation.
- Verify the deployment certificate signature, expiry, broker, account, policy and capital.
- Confirm the kill switch is inactive.
- Confirm broker authentication and account reachability.
- Confirm the correct paper or live endpoint.
- Confirm market-data freshness and secondary-feed agreement.
- Verify host time synchronization.
- Verify audit database writeability and hash-chain integrity.
- Run reconciliation against the broker.
- Verify monitoring and alert delivery.
- Confirm the daily change ticket and named operators.
During the session¶
Monitor at least:
- process heartbeat;
- broker latency and request failures;
- market-data age;
- signal age;
- order-entry rate;
- open orders;
- rejects and cancellations;
- fills and partial fills;
- positions and leverage;
- cash and buying power;
- realized and unrealized PnL;
- daily loss and drawdown;
- reconciliation status;
- audit-chain validity.
Automatic stop conditions¶
- daily loss limit;
- drawdown limit;
- repeated broker failures;
- reconciliation mismatch;
- stale or invalid market data;
- account block;
- unauthorized symbol;
- price-collar violation;
- invalid certificate or expired certificate;
- corrupted kill-switch or audit state.
Manual stop procedure¶
- Call
engine.emergency_stop(reason, operator=...). - Confirm open-order cancellation results.
- Verify broker positions directly in the broker interface.
- Flatten positions manually only under the documented emergency policy.
- Snapshot and back up the audit database.
- Open an incident ticket.
- Preserve logs and market-data evidence.
- Do not clear the kill switch until root-cause analysis and approval are complete.
Restart procedure¶
- Keep the kill switch active.
- Restore the latest verified audit backup if necessary.
- Verify the hash chain.
- Query broker account, positions and open orders.
- Reconcile expected and actual state.
- Resolve every mismatch.
- Repeat readiness checks affected by the incident.
- Issue a new deployment certificate if release, environment, risk policy, account or capital changed.
- Clear the kill switch with an accountable operator.
- Resume first in shadow or paper mode when practical.
End of day¶
- Reconcile positions, cash, orders and fills.
- Verify the audit hash chain.
- Create an immutable backup.
- Generate daily risk and execution reports.
- Record incidents, rejects and limit events.
- Revoke or allow the short-lived certificate to expire.
- Review whether capital limits remain appropriate.