Based is an AI foundry focused on quantization, fine-tuning, and inference infrastructure for frontier-scale systems. We make large models practical — reproducible, fast, and deployable on real hardware.
From raw checkpoints to production serving. We handle the hard parts of working with large open-weight models.
W4/W8 and sub-4-bit (IQ1S) quantization of large models. Reproducible recipes that preserve quality while fitting on real hardware.
LoRA and overlay fine-tuning of frontier open-weight models — Kimi, GLM, and others — for specific domains and tasks.
Production inference on multi-node clusters with tensor parallelism (TP4), long context (200K), and low-latency decode. We publish the exact serving recipes.
Operating DGX Spark and GB10-class hardware. Storage, orchestration, and reproducible builds across multi-node deployments.
We publish our work openly — checkpoints, buckets, and recipes on Hugging Face, with full build scripts on GitHub.
Hands-on guidance for teams adopting open-weight models — hardware selection, quantization strategy, and serving architecture.
Everything we build is published. Models and checkpoints on Hugging Face, build and serving recipes on GitHub.
Tell us what you're working on. We'll get back to you within one business day.