lab358, Inc.

Stealth

About

lab358 hosts open-weight language models with a retrieval-attention serving path. Instead of pasting documents into every prompt, lab358 precomputes each document once into a key-value cache, stores it durably, and attends over it at inference time. Documents are precomputed once, stored durably, and can be structured to greatly reduce the cost of multiple documents in the context, with the live sequence bounded by the GPU. The platform either runs as a cloud service or deploys into the customer's own cloud account, so weights, documents and inference stay inside their boundary. It ships as a Helm chart with an OpenAI-compatible API, a console, and a management CLI.

This startup is in stealth mode

The founder is keeping the full profile private for now — only the company name and description are public. Request access and, once the founder approves, you'll see the complete profile including traction, team, products, and updates.