03 / Run it yourself

Control where the work runs.

Your infrastructure. Your model choices. An explicit boundary for anything that leaves it.

Inside your boundary

Local keeps authority.

Merlin can use a local inference fleet while the runtime governs tools, workspace changes and final answers. Model routing and permission to act are separate responsibilities.

At the boundary

Cloud is a configuration choice.

Cloud escalation exists for long context and other configured routing policies. It can be switched off. If enabled, requests sent to those providers can carry context; review the actual profile before treating a workload as local-only.

A request enters Merlin's plan, analyze, execute and verify workflow. Evidence and context, model routing, and tools and workspace connect to the orchestrator. Model routing connects to a GPU inference fleet and optional cloud fallback.
Conceptual architecture · not every request follows every stage. GPU symbols are representative.View full size ↗

Deployment considerations

Choose the workload before the hardware.

Model size, quantization, context length and simultaneous requests all affect capacity. A diagram with three accelerator symbols is not a three-GPU requirement, and there is no single memory figure that fits every deployment.

Before a first run

Make these decisions explicit.

Which models may receive the workload? Which sources and tools may it access? Which provider fallbacks are enabled? What evidence is required before changes and completion?

From the lab

The machines behind the work.

The fleet deserves to be shown as it is. The photograph is still to be captured; the architecture above is a conceptual drawing.

Inspect the contract

Start with the requirements.