Where the Computation Lives
The question is not which model runs. It is whose machine it runs on, and what that ownership changes.
TAM-CMN.08 · The Common Mind · The Approximate Mind
There is a server room in Bhubaneswar that used to be a storage closet. It is small, cooled by two wall-mounted units that struggle in April, and it holds four machines that together cost less than a mid-range car. The district health authority bought them with a budget line that was meant for office furniture. A systems administrator named Priya, who is twenty-six and was hired to maintain spreadsheets, keeps them running. She has taught herself more in two years than her computer science degree covered in four, mostly because the machines break in ways the textbooks did not anticipate, and she is the only person who can fix them.
The models that run on Priya’s machines serve eleven district clinics. They handle drug interactions, flag abnormal lab trajectories, and assist with triaging the queue of patients who arrive each morning before the doctors do. The models are not large. They were distilled from bigger ones, compressed and fine-tuned for the specific work of a district health system in Odisha, and they run inference locally, which means that when a nurse in a village clinic asks a question, the question does not leave the state. The answer comes from a machine in Bhubaneswar, in a room that used to hold filing cabinets, maintained by a woman who was not trained for this job and has become indispensable to it.
This is what local inference looks like when it is real rather than theoretical, and it is both more possible and more fragile than most arguments about decentralization acknowledge.
The possibility is genuine. Inference, the act of running a trained model to produce an answer, is far cheaper than training. The hardware it requires is modest and getting cheaper. A machine that can run a capable domain model costs what a district can budget for, not what a country must allocate. The models themselves, when open-weight, can be downloaded, fine-tuned, and deployed without asking permission from anyone, without sending data to a foreign server, without paying rent on every query. This is a real form of ownership in a domain where ownership is rare, and it changes the economics of who can provide AI capability from a question about billion-dollar labs to a question about storage closets and the people willing to maintain them.
Owning the inference is not the same as owning the intelligence. But it is the part of ownership that is within reach.
The fragility is also genuine, and the series has to be honest about where ownership ends and dependency begins. The models running on Priya’s machines did not originate there. They were distilled from larger models, which means their capability flowed downhill from a frontier someone else built, using hardware someone else manufactured, trained on a corpus someone else assembled. Priya can run the models. She cannot retrain them from scratch. She cannot build the chips they run on. She cannot reproduce the research that produced the architecture. Local inference is sovereignty over the last mile, and the last mile matters, but it is not sovereignty over the road.
This dependency has layers, and the layers matter because they determine what can be changed and what cannot. The hardware layer is the most concentrated. The chips that train and run the most capable models are designed by a small number of companies and manufactured in a smaller number of fabrication facilities, and no amount of local ownership at the inference layer changes that concentration. A country that runs its own models on imported hardware is less dependent than a country that rents inference from a foreign cloud, but it is not independent, and a supply disruption at the chip level reaches everyone downstream.
The model layer is less concentrated than the hardware layer, and this is where open weights change the equation. A model released with open weights can be copied, modified, fine-tuned, and deployed by anyone with the hardware to run it. The distillation that produced the models on Priya’s machines was possible because the parent model’s weights were open. Had they been closed, had the capability been available only as a rented service through an API, the district health authority would be paying per query to a company that could change the price, change the terms, or shut off access at any time. Open weights do not eliminate the dependency on the original training run. But they convert a rental into a possession, and a possession can be maintained, modified, and passed on in ways a rental never can.
The difference between open and closed at this layer is the difference between a book you own and a book you are permitted to read in someone else’s library. The library can close. The book on your shelf cannot be taken back. For Priya’s machines and the clinics they serve, this distinction is not philosophical. It is the difference between a system the district controls and a system the district borrows.
And the book, once owned, can be annotated. This is what fine-tuning means in practice, and it is where local ownership begins to produce local intelligence rather than merely hosting foreign intelligence locally. Priya’s team took the open-weight model and trained it further on the district’s own health data: the patterns of disease that are specific to this region, the drug availability that is specific to these pharmacies, the seasonal rhythms that are specific to this population. The model they ended with is not the model they started with. It has been shaped by the place it serves, and that shaping is something the original builders could not have done because they did not have the data and would not have had the reason. The fine-tuned model is a hybrid: its deep architecture came from elsewhere, but its surface knowledge is indigenous, and the surface is where the patient meets it.
This is also where capability begins, tentatively, to flow upward rather than only down. A model fine-tuned on local disease patterns may discover correlations the parent model never saw, because the parent model never looked at this population in this granularity. A district that aggregates its fine-tuning insights and shares them with other districts, or with the national health authority, is contributing knowledge back into a system rather than only consuming from it. The flow is small now. Whether it can become structural, whether the periphery can generate intelligence that the center needs, is a question the series returns to.
Training is where the dependency runs deepest, and this is the layer the floor cannot fully resolve in its current form. Training a frontier model requires hardware at a scale only a few organizations possess, data at a scale only a few can assemble, and expertise at a concentration only a few locations in the world maintain. A district health authority in Odisha will not train a frontier model. A country the size of India might, and the series has already argued that India’s capacity to do so is a legitimate answer to the sovereignty question. But most countries cannot, and for them the dependency at the training layer is structural, not temporary. The floor does not pretend otherwise.
What the floor can do, and what matters for the people it serves, is concentrate ownership where ownership is achievable and honest about where it is not. Local inference on open-weight models, fine-tuned on local data, is achievable sovereignty. It means the data stays local, the capability is available regardless of what a foreign company decides, and the cost is a capital expense rather than an ongoing rent. The training dependency remains, but it is a dependency on a thing that has already been released into the world, not on a service that can be withdrawn.
I wonder whether the most durable kind of sovereignty is not the kind that eliminates every dependency but the kind that knows exactly which dependencies it has chosen and can name what it would take to dissolve each one.
Priya does not think about any of this in these terms. She thinks about cooling, about the UPS battery that needs replacing, about the model update she tested last weekend that broke the lab-flagging module until she rolled it back at two in the morning. She thinks about the fact that the nearest person who understands what she does is in Hyderabad, eight hours by train, and they have never met in person. She is building something that looks, from the outside, like a minor IT installation in a government office. From the inside it is the beginning of a district that owns its own cognitive infrastructure, maintained by someone who was not supposed to be the person who does this and has become, by necessity and stubbornness, exactly that person.
The storage closet is small and hot and the wall units struggle in April. The machines run anyway. The nurses ask their questions and the answers come back and the data does not leave the state. Priya checks the logs before she goes home, the way a farmer checks the field before dark: not because something is expected to go wrong, but because the thing is hers, and you watch what is yours.
How this essay connects to others across The Approximate Mind.
