OneBit

What we publish

Research

Ternary models, learned quantization boundaries, and what it takes to keep reasoning intact at 1.58 bits.

Papers

01 · CLOE V1.2

Post-Training Ternarization of Qwen3 Language Models

Rotation, ternarization and error compensation on Qwen3-4B, end to end: capability, effective bits per weight, storage and inference, measured.

Technical report · August 2026 · Malik, Devan, Mehra
PDF
02 · CLOE V1.1

Capability-Stratified Degradation in Ternary Language Models

What survives when a pretrained 752M model is pushed to three states: not a uniformly weaker model, but a stratified one.

Technical paper · 2026 · Malik, Devan, Mehra
PDF
03 · CLOE V1.0

Cloe: Hybrid Surgical Distillation for 1.58-Bit Edge Quantization of Large Language Models

Attention stays at 16 bits, the MLPs go ternary, and a teacher restores what the cut removed.

Paper · 2026 · Malik, Vishnuprasad, Devan, Mehra
PDF
04 · ITBO

A Dynamic Ternary Architecture for Latency-Critical Edge Agents

Learned quantization boundaries let a ternary network decide, per layer, where a weight becomes zero.

Paper · February 2026 · Malik, Mehra, Vishnuprasad, Pundir, Tyagi, Arutkeerthi
PDF
05 · Overview

Edge AI using Ultra Low Bit LLMs

Why the next generation of models will run on the chips that already exist, and what that changes.

White paper · January 2026 · Malik, Mehra, Vishnuprasad, Pundir, Tyagi
PDF

In one line

Every weight is −1, 0 or +1. The boundaries that decide which are learned, not fixed, so each layer chooses its own sparsity and the circuits that carry reasoning survive.

Full methodology, hardware and raw logs accompany every number we publish. A number without them does not ship.