Korrin

Core AI Inference Infrastructure.

Infrastructure for the inference economy.

Your data. Your models. Your intelligence.

Performance at a glance

600tok/s
Per RDU rack vs 125 on GPU
10×
Tokens per kW vs GPU infrastructure
55–70%
Lower cost per token vs GPU cloud
10 kW
Power draw per rack
Air Cooled
No liquid cooling required
1 rack
Per model instance

Models optimized to run natively on ASIC chips for maximum performance with minimal power

1
Models

Open-weights LLMs

2
ASIC Compiler

Compiled for native ASIC silicon

3
Inference Infra

Purpose-built hardware inference

4
Your Endpoint

One endpoint, multiple models & delivery methods for privacy & control

The result is up to 10× more tokens per kilowatt and 55–70% lower cost per token versus GPU cloud providers.

Choose the deployment model that fits your security, performance & budget

Lowest Cost
Shared Endpoint

Multi-tenant inference. Lowest cost entry point. Ideal for development and early production.

Performance
Dedicated Endpoint

Reserved capacity. Higher throughput and isolation. Production workloads.

Managed
Hosted Endpoint

Single-tenant managed environment. SCX operations. Enterprise grade.

Full Control
On-Prem Endpoint

Maximum control. Deploy in your own data center. Full data isolation.

Zero retention. Zero training. Full control.

No Prompt Caching

Queries are processed and discarded. Never stored or indexed for future use.

No Input/Output Retention

Responses delivered to you only. No copies kept on Korrin systems.

No Training on Your Data

Your proprietary information is never used to improve or train any model.

No Resale or Reuse

Your data, insights, and competitive advantage are never monetized.

Access every model through a single OpenAI-compatible API.

  • OpenAI-compatible API
    Change one line of code. Same SDK, same format, same tools.
  • Full model catalogue
    Llama 4, DeepSeek V3, MiniMax M2.7, Qwen 3.5, and more.
  • 192K context windows
    Run large context workloads with guaranteed data privacy.
  • Tool calling and RAG
    Native function calling, embeddings, and retrieval pipelines.
  • Fine-tuning
    Bring your own weights. Private hosting with full IP control.

Leading open-weights models compiled & optimized

Live on SCX Node 1
GPT-OSS-120B & 20BMeta Llama 3 & 4DeepSeek V3.1OpenAI WhisperQwen 3Qwen 3-TTSMistral-E5
Coming This Month
Gemma 4DeepSeek V4Mistral 3MiniMax M2.7

Leading open weights models perform many tasks better

Chat & Language
gpt-oss-120b, DeepSeek V3.1 Terminus, MiniMax M2.7, Llama 4 Maverick, Qwen 3.5, Gemma 4
Image, Vision & Multimodal
Llama 4 Maverick, Gemma 4
Audio & Voice
Qwen 3
Transcription
OpenAI Whisper
Code Generation
SCX-Coder (192K context), DeepSeek V4 Pro
Embeddings & Search
E5-Mistral-7B
Rerank Engine
MiniMax M3
Moderation & Guardrails
Llama 4
Localization
Korrin MAGPiE