Where the Computation Lives — Summary
There is a server room in Bhubaneswar that used to be a storage closet. Four machines, cooled by two struggling wall units, serve eleven district clinics. Priya, twenty-six, hired to maintain spreadsheets, keeps them running. The models handle drug interactions and lab trajectories, and when a nurse asks a question, the question does not leave the state.
Local inference is both more possible and more fragile than most arguments acknowledge. The hardware is modest and getting cheaper. Open-weight models can be downloaded, fine-tuned, and deployed without permission or rent. This is real ownership. But the models were distilled from larger ones, built on hardware someone else manufactured, trained on a corpus someone else assembled. Priya can run them. She cannot retrain them from scratch. Local inference is sovereignty over the last mile. It is not sovereignty over the road.
The layers of dependency matter. Hardware is most concentrated: chips designed by a few companies, manufactured in fewer facilities. The model layer is less concentrated, and open weights change the equation, converting a rental into a possession. A possession can be maintained and modified. A rental can be withdrawn. Fine-tuning is where ownership starts producing local intelligence: Priya’s team trained the model on the district’s own health data, and the resulting model carries indigenous surface knowledge even though its deep architecture came from elsewhere. Capability begins, tentatively, to flow upward.
Priya checks the logs before she goes home, the way a farmer checks the field before dark. Not because something is expected to go wrong, but because the thing is hers, and you watch what is yours.