System v1.0 Online

Local AI. Infinite Scale.

A highly-scalable distributed inference architecture. It features a high-throughput Go gateway orchestrating real-time SSE streams from decoupled PyTorch gRPC workers, paired with Supabase for seamless edge-to-cloud state synchronization.

grpc_stream.log

gRPC Microservices

High-performance backend. Real-time Server-Sent Events (SSE) stream tokens from a Go Gateway, backed by raw gRPC PyTorch workers.

Hybrid Architecture

Heavy inference runs locally on isolated Docker nodes, while your chat sessions securely sync across devices via Supabase.

Open Weights

Swap between models like Gemma 3, SmolLM2, and TinyLlama seamlessly. The routing mesh handles the inference pipeline automatically.

Built for Scale

The Inference Mesh architecture routes your requests instantly from browser to GPU.

Client
Go GatewaySSE / gRPC Bridge
Python WorkersPyTorch / HuggingFace

Supported Models

Hot-swap between cutting-edge open-weight models instantly.

Gemma 3

Google DeepMind

Parameters270M
Context Window8k Tokens
Avg. ThroughputUp to 80 tk/s

SmolLM2

HuggingFace

Parameters1.7B
Context Window8k Tokens
Avg. ThroughputUp to 60 tk/s

TinyLlama

TinyLlama Project

Parameters1.1B
Context Window2k Tokens
Avg. ThroughputUp to 60 tk/s

* Maximum throughput requires dedicated GPU hardware (CUDA/MPS). CPU fallback will be significantly slower.