Quickstart

TensorGrid exposes an OpenAI-compatible REST API and a gRPC streaming interface.

Authentication

All requests require an API key in the Authorization header.

curl https://gpu.execut4ble.site/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"llama-3.3-70b-instruct","stream":true,"messages":[{"role":"user","content":"Hello"}]}'

Endpoints

MethodPathDescription
POST/v1/chat/completionsChat completions (SSE streaming)
POST/v1/embeddingsText embeddings
GET/v1/modelsList available models
gRPCtensorgrid.compute.v1.ChannelBidirectional token streaming

Rate limits

Limits are enforced per API key. Exceeding them returns 429 with a Retry-After header.