Now serving inference from UK GPU clusters

Frontier inference
at gigascale.

Gigatokens delivers frontier inference to financial institutions and enterprises across the globe — built by Behavox, trusted by the world’s most regulated firms. Now open to every team that needs it. API access to the best open-weights models — plus Behavox-built models optimised for compliance — served from dedicated GPU clusters in the United Kingdom.

status.gigatokens.ai GT-LDN-1 · OPERATIONAL
  • kimi-k3172.4k tok/s
  • glm-5.3118.6k tok/s
  • glm-5.3-flash137.9k tok/s
  • deepseek-v4-pro78.2k tok/s
  • qwen3.8-max51.3k tok/s
  • behavox-quantum38.4k tok/s
TOKENS TODAY 48,236,521,326
50B+
tokens served daily
<180ms
median time to first token
99.95%
uptime SLA
🇬🇧 UK
sovereign GPU infrastructure
Backed by
BlackRock SoftBank Citigroup
Infrastructure partners
Google Cloud Civo
What we sell

Four ways to buy tokens

Direct API for builders, aggregator listings for reach, Behavox Compliance products for regulated firms — and native access inside Gigatokens Conductor.

Direct API access

OpenAI-compatible endpoints for the leading open-weights models. Transparent per-million-token pricing, no rate-limit games, no cold starts.

Via OpenRouter & aggregators

Already routing through OpenRouter or similar platforms? Gigatokens endpoints are listed there too — same hardware, same latency, your existing billing.

Available on OpenRouter

Via Behavox Compliance products

Gigatokens power the Behavox product suiteQuantum communication surveillance, Polaris trade surveillance, Falcon insider-threat detection and Digital Employees — and the same Behavox models are available per token on the direct API.

Behavox product suite

Via Gigatokens Conductor

Gigatokens are available natively inside Gigatokens Conductor — the secure platform for running AI agents that Behavox uses internally. Agents and automations buy and spend tokens directly.

Gigatokens Conductor
Model catalogue

Serious models, honest pricing

A curated catalogue — every model runs on dedicated GPU clusters, benchmarked and load-tested before it ships. Prices per million tokens.

ModelTypeInput / MCached input / MOutput / M
Kimi K3Open weights$2.40$0.20$12.00
GLM 5.3Open weights$1.10$0.20$3.50
GLM 5.3 FlashOpen weights$0.06$0.02$0.20
DeepSeek V4 ProOpen weights$0.80$0.06$1.65
DeepSeek V4 FlashOpen weights$0.07$0.016$0.14
Qwen 3.8 MaxOpen weights$1.60$0.20$4.80
Behavox QuantumCommunication surveillance across 150+ channels and 50+ languages — voice, chat, WhatsApp and Teams, at the lowest alert volume.Compliance$5.00$0.50$25.00
Behavox PolarisCross-asset trade surveillance across all 10 asset classes — 35 regulation-mapped policies and an AI filter that cuts alert volume.Compliance$5.00$0.50$25.00
Gigatokens Conductor

Every agent gets its own computer

Conductor runs each AI agent on its own virtual machine — a real computer in the cloud, not a chat window. Close your laptop; your agents keep working. Behavox runs its own AI workforce this way — over half a million agent machines to date — and now offers it to Gigatokens customers with inference built in.

Isolation

A separate computer for every agent

The boundary between agents is a machine boundary — enforced by infrastructure, not by prompts.

  • Own files, memory and software — install anything, break nothing
  • Not tied to your laptop — agents run 24/7 in the cloud
  • Blast radius of one — freeze a machine, a group, or the whole fleet in one click
  • Scaling means adding computers — run one agent or a thousand
  • Stormgate — the only door out

    One connection out of every machine. Rules you set, read-only by default, full audit trail. Credentials injected per request — agents never hold a secret.

  • Replicas of your systems

    Agents rebuild your production environment — even on-prem — on machines of their own, and test every change there before anything real is touched.

  • Every machine accounted for

    Every token attributed to an agent, task and team. Budgets, caps, live cost telemetry — one console for the whole fleet.

  • Agent factories

    Pipelines of agents — each step on its own machine — that plan, build, review and deploy autonomously. Start from a library of proven templates and adapt them to your systems.

Want to know more about Gigatokens Conductor? Contact sales.

Infrastructure

Dedicated GPUs.
UK sovereign data processing.

  • Dedicated GPU capacity Dedicated clusters, predictable capacity, published latency — operated with proven infrastructure partners.
  • UK sovereign data processing Prompts and completions are processed in the United Kingdom, full stop. Built for firms that answer to regulators.
  • Zero retention by default We don't train on your data. Ephemeral processing with optional logging you control.
  • Enterprise-grade security Encryption in transit and at rest, SSO/SAML, audit logs, and dedicated-capacity options.
🇬🇧

London Region · GT-LDN-1

High-density GPU clusters purpose-built for large-model inference, with room to grow as demand does.

ISO 27001 (Civo)GDPR / UK DPAZero data retention

Start shipping on Gigatokens

Per-token pricing, dedicated capacity, enterprise onboarding — tell us what you're building.

Contact sales

We use what you send to answer your enquiry. See our privacy notice.

Browse models