How local people train the model
Foundation models learn from text on the internet. The people who know most about water in the Bengal Basin rarely write there: they speak Bangla and its dialects, they keep records on paper, and many of them are young. The Blueprint's model is trained by them, on data they consent to share, and judged by them before it is used.
Who trains it
Community-based organisationsGroups rooted in their own villages and neighbourhoods that have worked on arsenic, salinity and safe water in the region for decades bring what no dataset holds: well-testing records, field reports and hard-won judgement about what fails and why. Their members also check every label before it reaches the model.
HouseholdsFamilies who drink the water grade it from A to F and leave voice notes in their own words and dialect about taste, colour, queues, breakdowns and cost. These are the voices the grade is built on.
YouthYoung people aged 18 to 35, who are rarely represented in foundation model training data, train as certified Watergraders. They collect reports, transcribe and translate dialect into standard Bangla, label the data, and are paid for it.
OperatorsThe women and men who run Water ATMs every day log faults, filter regenerations and water tests. That turns daily operating experience into training data.
Six steps from voice to model
- Consent first. Every contributor agrees to how their data is used and can withdraw it. The data sits in Bangladesh under a community data trust, governed by the city, community-based organisations, youth representatives, and Drinkwell. The data trust follows principles drawn from Indigenous data sovereignty, the CARE Principles: collective benefit, authority to control, responsibility and ethics.
- Collect in local languages. Households and operators contribute voice notes, photos and grades through WaterGrade and the Water ATMs, in Bangla and its dialects. Community-based organisations contribute their records, with permission.
- Label with two keys. Youth Watergraders transcribe, translate and label. Members of community-based organisations validate. A record only enters the training set when both have signed off.
- Fine-tune a small open model. We start from an open model that already reads Bengali, then fine-tune it on this data. It is small enough to run on modest hardware in Bangladesh. The weights are released under an open licence, with a model card that credits every contributing group.
- Community evaluation. Community-based organisations, youth and households write the test questions and score the answers. The model is only used once it passes their benchmark, not just ours.
- Trace every decision. Each training record carries its provenance: who, when and under what consent. Every model output that triggers an action, such as a repair, a filter regeneration or a water test, is logged to the public ledger, where households' grades show whether it was right.
Guardrails: the model advises and people decide. It never declares a source safe to drink without a laboratory test, and it holds no names or personal details. Contributors are credited, and part of the Blueprint's funding pays for their time.