Compare

Most agent tools watch the work. Alchemy governs it.

Observability shows you what happened. Frameworks leave governance to you. Hosted runtimes keep the loop on someone else's infrastructure. Alchemy runs an organization of agents on yours, with spend ceilings, approvals, an independent check against a definition of done, and a record you can verify.

Self-hostedEvery step governedNothing metered
01 / Capability matrix

Side by side. What each category gives you out of the box.

The other columns describe each category in general terms. Individual products vary, and many teams use more than one.

CapabilityAlchemyObservabilityOrchestration frameworksHosted runtimes
Spend ceilingsPer task, plus daily and monthly limits per project and account, enforced before each stepReports spend after the factBuilt by youBudgets on the vendor's platform
Approval of actions you markHeld for an administrator; no agent can approveNot in scopeMiddleware you add per toolPer-tool permission settings
Definition of doneEvery deliverable opened by an independent member before a run can finishScores traces after the runWhatever your code checksVaries by product
What the agent saw at each stepKept for every step, exactlyTraces sent by your instrumentationCheckpoints of workflow stateSession logs in the vendor's store
Outside actions after a crashReconciled, never repeatedNot in scopeYour code's responsibilityVaries by product
Who may touch whatPer-member reach and delegation in one manifestNot in scopeWired in your codePer-agent tools and policies
The recordEncrypted, hash-chained, exportable and independently verifiableSearchable tracesThe state store you configureLogs held by the vendor
Where it runsYour infrastructureVendor cloud or self-hostedWherever you deploy itThe vendor's infrastructure
PricingLicence per organization; nothing meteredUsually metered by trace, span or seatOpen-source libraries; hosted tiers meteredMetered by usage and runtime
02 / Versus observability

A trace tells you what happened. Alchemy decides what may happen next.

Tools such as LangSmith, Langfuse and Datadog capture prompts, outputs and costs, and they do it well. But a trace is assembled after the call, from what your code chose to send.

Alchemy keeps its own record as the work happens: exactly what each agent was given, and every action it took on your systems. Spend is checked before the next step, not charted after it. The record stays encrypted on your disk, hash-chained, and can be exported and verified without us. Keep your dashboards; Alchemy adds the controls they were never built to enforce.

03 / Versus orchestration frameworks

Frameworks give you building blocks. Governance is still yours to build.

LangGraph, CrewAI and similar frameworks are flexible ways to wire agents together. Approvals, spend limits, access boundaries and an audit trail are left to your code, one tool and one project at a time.

In Alchemy they are declared once, in the manifest. It names every member, what each may touch and to whom it may hand work, which actions wait for an administrator, and what counts as done. The engine enforces all of it at every step, and the manifest is one file your security team can review.

04 / Versus durable execution

Durable engines make steps reliable. They do not know what your agents are for.

Temporal, Inngest and similar engines survive crashes and retry steps well. They bring no organization of members, no approvals declared per action, no spend ceilings on model use and no definition of done; you build those on top.

Alchemy is that layer, built for agents. A crash costs only the steps in flight, and an outside action is never repeated after one: it is reconciled against what actually happened, and anything uncertain goes to a person.

05 / Versus hosted runtimes

Hosted runtimes run your agents for you. Alchemy keeps them with you.

Model and cloud providers now offer hosted agent runtimes with budgets, tool permissions and evaluation. They are convenient, and the agent loop and its record live on the vendor's infrastructure.

Alchemy runs the loop, the approvals and the record on infrastructure you control, with your own Anthropic API key. Nothing is sent to Agiliti, model spend is never marked up, and the evidence of every run stays on your disk, in a form an auditor can verify independently.

06 / Versus control planes and gateways

Control planes govern the boundary. Alchemy governs the work inside the run.

Enterprise agent registries, AI gateways and policy layers inventory agents, route requests and block calls a policy forbids. From the edge, they see a request.

Alchemy sees the run: which member asked, what it had been handed, whether the action needed approval, and whether the deliverable was checked before the work was called done. The two layers are complementary. A gateway decides which requests may leave your network; Alchemy decides which work may proceed, who signs off on it, and what evidence it leaves behind.

07 / When Alchemy fits

When Alchemy fits. Agent work that has to answer to someone.

Agents act beyond your walls

They push code, open pull requests, file, publish or pay for services, and those actions must wait for a named person.

Spend has to be predictable

Finance wants ceilings per task, project and month, and no metered invoice from the platform.

Evidence has to hold up

Audit and compliance need a record the model did not write, kept on your infrastructure and verifiable without the vendor.

Data has to stay in-house

Alchemy runs on your own Windows or Linux infrastructure, model use is billed to your own Anthropic key, and nothing is sent to Agiliti.

One team of agents, not a pile of scripts

You want members that plan, delegate and deliver as one organization, with independent tasks running side by side within limits you set.

Done has to mean done

Every deliverable is opened by an independent member against a written definition before a run can finish.

08 / FAQ

Questions buyers ask when comparing. Answered plainly.

How is Alchemy different from LangGraph or CrewAI?

LangGraph and CrewAI are frameworks for writing agent logic in code; governance controls such as approvals, budgets and audit are largely yours to assemble. Alchemy is a governed runtime: you describe the organization in a manifest, and it enforces approvals, spend ceilings, definition-of-done checks and a tamper-evident record while the work runs, on your infrastructure.

How does Alchemy compare with hosted agent runtimes?

Hosted runtimes run your agents on the vendor’s infrastructure and typically meter usage. Alchemy runs as one process on infrastructure you control, meters nothing and uses your own Anthropic API key, so records, deliverables and spend stay with you. The trade-off is that you operate it yourself instead of renting a managed service.

Is Alchemy an observability tool?

No. Observability tools show what agents did after the fact. Alchemy enforces rules while the work happens: it checks ceilings before each step, holds marked actions for a person’s approval and will not finish a run until an independent member has checked the deliverables. Its tamper-evident record comes from that enforcement, not from sampling traces.

Can I use Alchemy with models other than Claude?

No. Alchemy runs Claude models and is built on Anthropic’s Claude Agent SDK. If you need several model providers in one system, a multi-provider framework may suit you better. If Claude fits your work, Alchemy gives you a governed runtime with approvals, spend ceilings, verification and a complete record without building those controls yourself.

When is a framework a better choice than Alchemy?

Choose a framework if you are embedding a single agent in your own product, need models from several providers, or want to hand-code agent control flow. Choose Alchemy when a team of agents does consequential work and you need approvals, spend ceilings, independent verification and an audit-ready record running on your infrastructure from day one.

Governed agent work, on your infrastructure. Nothing metered.

Start with an evaluation, or build your first organization of agents with us in an eight-week design partnership.