An OpenAI-compatible LLM access point.
loch-nessh is a VRAM-aware model broker — a standard, OpenAI-compatible chat/completions/embeddings/images endpoint in front of a real GPU fleet, with queueing, scheduling, and node failover handled for you. Point any OpenAI-compatible client at it and go.
Live status →
Cluster capacity and aggregate usage, publicly visible.
Authorized dashboard →
Full queue/audit/billing view — requires access.
API surface
POST /v1/chat/completions
POST /v1/completions
POST /v1/embeddings
POST /v1/images/generations
POST /v1/audio/transcriptions
GET /v1/models