Once Arceo has watched enough live calls, the monthly number lands inside this range.
Sequences of safe-looking actions that turn dangerous in order.
Already classified across 11 services. New actions are classified automatically.
Pennies a call. Some of them you can't take back.
Every call your agent makes costs a few cents and carries a weight. Reading barely counts. Anything you cannot undo counts double.
One score for how much damage an agent could do.
Every action carries a weight. Add them up, double anything you cannot undo, and you get a number out of 100.
Anything a delete or a payment touches counts double, because you cannot take it back. Above 60, an agent needs a policy before it ships.
Two safe agents. One dangerous chain.
Agents hand work to each other. Arceo follows the handoff, so a sequence that is only dangerous across two agents still gets caught.
Cost and risk, in one report your CFO will sign.
The forecast tells you how much to trust it
Day one you get a number and a wide range. Watch a week of real traffic and the range closes to ±15%.
Know which lever actually moves the bill
Arceo nudges each input and ranks what moves. Call volume wins by a distance; a daily call cap is the control that holds a budget.
Call volume swamps everything else. Cap it and you have capped the bill.
Catch the risks that only show up in sequence
Reading a customer record is fine. Sending an email is fine. Doing both in a row is a data leak. Arceo watches for 32 of these pairs across 10 kinds of risk.
Each red cell is a sequence that has already gone wrong at a real company. Read a record then email it out, and you have the shape of the Copilot data leak.
The check that stops it before it merges.
Arceo runs as a GitHub Action. It scores every agent in the diff, posts the report as a comment, and fails the build when one crosses the threshold you set.
Everyone else measures agents after you deploy them
Observability tells you what you already spent. Security tooling tells you what already broke. Arceo answers the question that gates deployment, before the agent goes live.
- Cost and risk for AI agents, in one report a finance team can sign off on
- Pre-deployment: the answer arrives before the agent handles a real request
- Platform-agnostic: Anthropic, OpenAI, MCP, GitHub, LangChain, or your own code
- Evaluation platforms score answer quality; Arceo prices what the agent can reach
- Security tools sell to the CISO; Arceo reports to the CIO and the CFO
- Observability measures spend after deploy; Arceo forecasts it before
Arceo governs the agents you build.
Three things you can check before you trust the number
We are early. Every claim below is checkable against the product.
The cost model is backtested
We re-priced 829 real Anthropic usage records, captured independently of the forecaster, through the engine; the high-confidence tier reproduces the actual spend. The test runs in CI, so a change that breaks pricing fails the build.
The risk model is mapped to published frameworks
Chain detection runs on risk-label transitions, so it generalises across every tool and vendor. The rule set maps to OWASP Agentic Security categories and MITRE ATT&CK tactics: privilege escalation, credential access, defense evasion, and collection.
The backend has been independently audited
A full security audit covering authentication, tenant isolation, injection, cryptography, dependencies, logging, and cost abuse returned zero critical findings. We share the report, and current remediation status, under NDA.
Read the security page