New · batch embeddings up to 10 Gbit/s

On-demand GPU compute and model serving.
Built for throughput.

TensorGrid gives you a single OpenAI-compatible endpoint for the best open-weight models. Stream tokens over HTTP/2 or gRPC, run bulk embedding jobs, and pay only for what you use.

Get startedView API reference

OpenAI-compatible

Drop-in replacement. Change the base URL and keep your existing SDKs, tools and prompts.

Streaming first

Server-sent events and bidirectional gRPC streams with sub-100 ms time-to-first-token.

Bulk & batch

Embeddings, re-ranking and offline batch jobs on dedicated high-bandwidth GPU pools.

Private by default

No prompt logging, no training on your data. Regional routing and per-key quotas.

Models

ModelContextTypeLatency (p50)
llama-3.3-70b-instruct128kChat78 ms
qwen2.5-72b-instruct128kChat84 ms
mistral-large-2128kChat91 ms
deepseek-v364kChat / Code103 ms
bge-m38kEmbeddings12 ms
whisper-large-v3-Speech-to-text240 ms

Pricing

Developer

$0 / month
1M tokens included, 10 req/s

Scale

$0.40 / 1M tokens
Priority routing, 200 req/s, SSE + gRPC

Dedicated

Custom
Reserved GPUs, 10 Gbit/s ingest, private regions

Streaming over gRPC

For long-lived sessions and bulk pipelines use our bidirectional gRPC endpoint. One connection, many concurrent streams.

# grpc endpoint
host     = gpu.execut4ble.site:443
service  = tensorgrid.compute.v1.Channel
metadata = authorization: Bearer <API_KEY>