Features

Every knob, proven.

The full feature inventory. Every claim below is claim safe and backed by the build.

Four engines, one app.

vLLM, llama.cpp, MLX, TensorRT-LLM. Launch, configure, and manage instances without touching a terminal.

Why it matters: Four engines, four flag grammars, four sets of docs. One app, one interface, zero terminal.

Hardware aware from second one.

Detects GPU, CPU, RAM including Apple unified memory. Computes real VRAM budgets.

Why it matters: No more out of memory roulette. You know what fits before you launch.

Tuning recipes.

Save, load, import, export, share. Community recipes built in. Hardware aware recommendations.

Why it matters: A proven recipe for your hardware, not a guess from a Discord thread.

Real benchmarking.

Throughput benchmark lanes with statistical safeguards. Honest numbers out of every run.

Why it matters: A lucky run is not an optimization. Statistical safeguards mean the number is real.

Quality lanes.

Long context needle in a haystack, instruction following, coding, math. Verifies the model still thinks, not just talks.

Why it matters: Faster and dumber is not an optimization. Quality lanes catch the silent regression.

Preflight checks.

11 pre launch checks that catch real failure modes: context overflow, KV cache mis sizing, quant and VRAM mismatches.

Why it matters: Catch the failure before it costs you a session.

Minefield Guard.

Live guardrails distilled from a registry of 116 real model serving failure traps.

Why it matters: 116 real failures, turned into guardrails that run before you press start.

Network auto discovery.

SSH probe scan finds GPU machines on your LAN, classifies them, one click to manage.

Why it matters: Your fleet appears in one window. No manual config, no spreadsheets of IP addresses.

Command Center.

Real time fleet dashboard: metrics, sparklines that survive restart, live GPU stats from every machine.

Why it matters: One window for the whole fleet. Monitoring that does not forget yesterday.

Tokens per Watt.

Efficiency as a first class metric. Green tuning is fast tuning.

Why it matters: The metric nobody benchmarks and everyone should. Efficiency is not a side note, it is the optimization.

Model Catalog.

Curated model list with a fit check: will it run on this machine, at what quant, with how much headroom. One click from maybe to running.

Why it matters: Three clicks from "I heard about that model" to it actually running, tuned for your machine.

Cross hardware comparison.

A/B configs and machines, CSV export.

Why it matters: Compare honestly. Export the data. Decide with numbers, not vibes.

Privacy and license.

Offline Ed25519 verified license. No account, no phone home, zero telemetry. Works fully offline.

Why it matters: Your models and your measurements never leave your machines. The license key is verified entirely offline.

Building now

Building now

  • Building now Agent Control Layer. An MCP server built into the app so AI agents can operate Orion Tuner natively. Localhost only, token authenticated, scoped, audited: every agent request is recorded in an audit trail. Your agent can discover hardware, pick or search models, build recipes, launch, benchmark, and iterate with the same guardrails as a human. Destructive actions sit behind an approval prompt. Your override always wins.
  • Building now Agent skill package. One click and your agent permanently knows every feature, knob, and workflow.
  • Building now CLI. Headless, same powers, scriptable, JSON output, for automation and agents without MCP.
  • Building nowSpeculative decoding A/B
  • Building nowSGLang engine support
  • Building nowMulti node mesh recipes

All items carry the "Building now" badge. No dates. No superlatives.

Orion Tuner

See what $10 buys.

Every feature above, $10 once, no subscription, offline and private.

sales@theorioncompanies.com