Side by side. What each category gives you out of the box.
The other columns describe each category in general terms. Individual products vary, and many teams use more than one.
| Capability | Alchemy | Observability | Orchestration frameworks | Hosted runtimes |
|---|---|---|---|---|
| Spend ceilings | Per task, plus daily and monthly limits per project and account, enforced before each step | Reports spend after the fact | Built by you | Budgets on the vendor's platform |
| Approval of actions you mark | Held for an administrator; no agent can approve | Not in scope | Middleware you add per tool | Per-tool permission settings |
| Definition of done | Every deliverable opened by an independent member before a run can finish | Scores traces after the run | Whatever your code checks | Varies by product |
| What the agent saw at each step | Kept for every step, exactly | Traces sent by your instrumentation | Checkpoints of workflow state | Session logs in the vendor's store |
| Outside actions after a crash | Reconciled, never repeated | Not in scope | Your code's responsibility | Varies by product |
| Who may touch what | Per-member reach and delegation in one manifest | Not in scope | Wired in your code | Per-agent tools and policies |
| The record | Encrypted, hash-chained, exportable and independently verifiable | Searchable traces | The state store you configure | Logs held by the vendor |
| Where it runs | Your infrastructure | Vendor cloud or self-hosted | Wherever you deploy it | The vendor's infrastructure |
| Pricing | Licence per organization; nothing metered | Usually metered by trace, span or seat | Open-source libraries; hosted tiers metered | Metered by usage and runtime |
A trace tells you what happened. Alchemy decides what may happen next.
Tools such as LangSmith, Langfuse and Datadog capture prompts, outputs and costs, and they do it well. But a trace is assembled after the call, from what your code chose to send.
Alchemy keeps its own record as the work happens: exactly what each agent was given, and every action it took on your systems. Spend is checked before the next step, not charted after it. The record stays encrypted on your disk, hash-chained, and can be exported and verified without us. Keep your dashboards; Alchemy adds the controls they were never built to enforce.
Frameworks give you building blocks. Governance is still yours to build.
LangGraph, CrewAI and similar frameworks are flexible ways to wire agents together. Approvals, spend limits, access boundaries and an audit trail are left to your code, one tool and one project at a time.
In Alchemy they are declared once, in the manifest. It names every member, what each may touch and to whom it may hand work, which actions wait for an administrator, and what counts as done. The engine enforces all of it at every step, and the manifest is one file your security team can review.
Durable engines make steps reliable. They do not know what your agents are for.
Temporal, Inngest and similar engines survive crashes and retry steps well. They bring no organization of members, no approvals declared per action, no spend ceilings on model use and no definition of done; you build those on top.
Alchemy is that layer, built for agents. A crash costs only the steps in flight, and an outside action is never repeated after one: it is reconciled against what actually happened, and anything uncertain goes to a person.
Hosted runtimes run your agents for you. Alchemy keeps them with you.
Model and cloud providers now offer hosted agent runtimes with budgets, tool permissions and evaluation. They are convenient, and the agent loop and its record live on the vendor's infrastructure.
Alchemy runs the loop, the approvals and the record on infrastructure you control, with your own Anthropic API key. Nothing is sent to Agiliti, model spend is never marked up, and the evidence of every run stays on your disk, in a form an auditor can verify independently.
Control planes govern the boundary. Alchemy governs the work inside the run.
Enterprise agent registries, AI gateways and policy layers inventory agents, route requests and block calls a policy forbids. From the edge, they see a request.
Alchemy sees the run: which member asked, what it had been handed, whether the action needed approval, and whether the deliverable was checked before the work was called done. The two layers are complementary. A gateway decides which requests may leave your network; Alchemy decides which work may proceed, who signs off on it, and what evidence it leaves behind.
When Alchemy fits. Agent work that has to answer to someone.
Agents act beyond your walls
They push code, open pull requests, file, publish or pay for services, and those actions must wait for a named person.
Spend has to be predictable
Finance wants ceilings per task, project and month, and no metered invoice from the platform.
Evidence has to hold up
Audit and compliance need a record the model did not write, kept on your infrastructure and verifiable without the vendor.
Data has to stay in-house
Alchemy runs on your own Windows or Linux infrastructure, model use is billed to your own Anthropic key, and nothing is sent to Agiliti.
One team of agents, not a pile of scripts
You want members that plan, delegate and deliver as one organization, with independent tasks running side by side within limits you set.
Done has to mean done
Every deliverable is opened by an independent member against a written definition before a run can finish.
Questions buyers ask when comparing. Answered plainly.
How is Alchemy different from LangGraph or CrewAI?
LangGraph and CrewAI are frameworks for writing agent logic in code; governance controls such as approvals, budgets and audit are largely yours to assemble. Alchemy is a governed runtime: you describe the organization in a manifest, and it enforces approvals, spend ceilings, definition-of-done checks and a tamper-evident record while the work runs, on your infrastructure.
How does Alchemy compare with hosted agent runtimes?
Hosted runtimes run your agents on the vendor’s infrastructure and typically meter usage. Alchemy runs as one process on infrastructure you control, meters nothing and uses your own Anthropic API key, so records, deliverables and spend stay with you. The trade-off is that you operate it yourself instead of renting a managed service.
Is Alchemy an observability tool?
No. Observability tools show what agents did after the fact. Alchemy enforces rules while the work happens: it checks ceilings before each step, holds marked actions for a person’s approval and will not finish a run until an independent member has checked the deliverables. Its tamper-evident record comes from that enforcement, not from sampling traces.
Can I use Alchemy with models other than Claude?
No. Alchemy runs Claude models and is built on Anthropic’s Claude Agent SDK. If you need several model providers in one system, a multi-provider framework may suit you better. If Claude fits your work, Alchemy gives you a governed runtime with approvals, spend ceilings, verification and a complete record without building those controls yourself.
When is a framework a better choice than Alchemy?
Choose a framework if you are embedding a single agent in your own product, need models from several providers, or want to hand-code agent control flow. Choose Alchemy when a team of agents does consequential work and you need approvals, spend ceilings, independent verification and an audit-ready record running on your infrastructure from day one.
Governed agent work, on your infrastructure. Nothing metered.
Start with an evaluation, or build your first organization of agents with us in an eight-week design partnership.
