The frontier is rented.
The edge is owned.
Edge is our model-efficiency and on-device research line. If the harness and the context are the moat, the engine underneath should be interchangeable, and eventually yours, running on hardware you already own. It is early, and we say so. This page explains what we are building and why.
We are not racing to the frontier.
We are racing to the edge.
The frontier race is a capital race: bigger clusters, bigger bills, a finish line that moves every quarter. For everyone not selling the compute, it is a race to the bottom. The race worth running is the other one: increasingly capable models on edge devices. On this view, India's AI story gets written by intrinsically capable hardware in ordinary hands, not by pay-as-you-go cloud.
For a law firm the argument is concrete. Privileged work can run on models that never leave the building. Frontier models stay within reach, and are reached for only when the work genuinely demands the reach. The firm sets each member's profile and stays in control of the spend. The engine becomes a choice the firm makes, matter by matter, not a dependency it inherits.
Four modes. One boundary.
Every piece of work runs in one of four confidentiality modes, ordered by who can see identifiable client data.
Inference runs on the device in your hand. The matter never crosses the air gap.
Models hosted inside the firm's own tenant. Data moves, but only to infrastructure the firm controls.
Identifying detail is stripped at the boundary. What crosses is the question, never the client.
The full reach of frontier models, reserved for work that carries no client confidence.
The dashed line is the privilege boundary. In Shielded, data crosses it only through the redaction gate; in Open, only work with nothing to protect crosses at all. The firm sets each member's mode profile and can read the meter.
The work is unglamorous, on purpose.
Four disciplines, each judged by one test: what it buys the firm.
Distillation
A large model teaches a small one the narrow work the firm actually needs. What it buys: capable behaviour at a size a device can hold.
Quantisation
The same model in a fraction of the memory, with the loss measured rather than guessed. What it buys: models that fit hardware the firm already owns.
Fine-tuning
Tuned on the firm's manner, never its matters: how the firm drafts, cites and reasons. What it buys: a small model that sounds like the firm from day one.
Evaluation
Indian legal benchmarks and long-context recall tests, run before anything is trusted. What it buys: knowing when a small model is enough, and when it is not.
Owned, swappable inference insulates a firm from vendor price rises, deprecations and outages. The cost is never being instantly on the absolute frontier. That is a trade, not a free lunch. For a law firm, it is the right trade.
Built for the device that exists, not the one we wish for.
The median Indian device is a 4 to 8 GB Android phone, not an M4 Mac. So the near-term shape is small models, roughly 1 to 4 billion parameters, doing the work they can do well, with a tiered bridge that overflows to the firm's own gateway when a task exceeds what the device can hold. Brief, and unglamorous, and honest about both.
Early, and worth watching.
Edge is research, not a product you can buy today. Leave your details and we will write when there is something real to show: a build, a benchmark, a device in hand. Nothing until then.