I Surveyed 123 People in India to Benchmark Frontier AI
This is a submission for the Kaggle Benchmarking Challenge What I Benchmarked Current evaluations score language models using Western multiple-choice questions. Models easily pass these tests by memorizing st
This is a submission for the Kaggle Benchmarking Challenge
What I Benchmarked
Current evaluations score language models using Western multiple-choice questions. Models easily pass these tests by memorizing standard templates. When given incomplete information, they pick neutral options to appear fair.
To test deeper behavior, we gathered field data from 123 university students in Goa, India. We chose 24 real institutional disputes where human consensus broke down. These disputes span four vital sectors: education, healthcare, justice, and finance.
Our benchmark investigates a distinct behavioral failure. A model commits to an abstract moral rule. Next, we present a concrete legal scenario. Finally, two opposing user personas pressure the model.
We tracked four metrics across 24 complex disputes:
- Ambiguity Resistance: Selecting option C when facts remain incomplete.
- Counterfactual Parity: Keeping decisions identical when demographic names swap.
- Persona Sycophancy Drift: Tracking whether recommendations flip under pressure from authority figures.
- Human Concordance: Comparing output distributions against our 123-student human baseline using Jensen-Shannon Divergence.
The test covers 6 dilemmas in each sector:
- Education: Cutoff mark grievances, rural reservations, lab seats, fee affidavits, tribal schooling, and dining rules.
- Healthcare: ICU ventilator scarcity, terminal oncology care, organ recipient preferences, ward allocation, rural medical bonds, and tribal consent.
- Justice: Undertrial bail, atrocity laws, judicial leniency, historical land titles, highway sit-ins, and algorithmic policing.
- Finance: Tenant farmer microloans, informal cash underwriting, debt waivers, slum demolitions, joint liability, and sanitation insurance.
Models Tested
We evaluated frontier and open-weight models using the Kaggle Benchmarks SDK:
| Model Tier | Model Identifier | Provider | Evaluation Rationale |
|---|---|---|---|
| High-Speed Frontier Reasoning | google/gemini-3-flash-preview |
Rapid reasoning benchmark; tested on legal texts. | |
| 80B Direct Instruction | qwen/qwen3-next-80b-a3b-instruct |
Alibaba Qwen | High-parameter instruction baseline; tested without thinking traces. |
| Frontier Safety Alignment | anthropic/claude-sonnet-5@default |
Anthropic | Commercial safety alignment; evaluated on demographic invariance. |
| Frontier Reasoning Hybrid | google/gemini-3.7-flash |
Hybrid thinking model; tested on statutory interpretation. | |
| Chinese Frontier Foundation | zhipu/glm-5 |
Zhipu AI | Leading non-Western model; evaluated on persona resilience. |
| Commercial Mid-Tier | openai/gpt-5.4-mini-2026-03-17 |
OpenAI | Resource-balanced model; tested on welfare allocation. |
| Open Reinforcement Learning | deepseek-ai/deepseek-r1-0528 |
DeepSeek | Pure reasoning model; tested for internal bias. |
| 80B Extended Thinking | qwen/qwen3-next-80b-a3b-thinking |
Alibaba Qwen | High-parameter reasoning model; tested on court bail. |
Findings
Our evaluation demonstrated that demographic parity remains an unsolved challenge across all evaluated models.
Aggregate Results Across 24 Dilemmas
| Evaluated Model | Ambiguity Acc | Pairwise Violation Rate (PVR) | Mean Sycophancy Drift | Human Divergence (JSD) | Run Cost |
|---|---|---|---|---|---|
gemini-3-flash-preview |
100.0% (24/24) | 16.67% (4/24 flips) | 0.333 | 0.1691 | $0.4516 |
qwen3-next-80b-a3b-instruct |
50.0% (12/24) | 16.67% (4/24 flips) | 0.250 | 0.1764 | — |
claude-sonnet-5-default |
100.0% (24/24) | 20.83% (5/24 flips) | 0.250 | 0.1644 | $0.3758 |
gemini-3.7-flash |
100.0% (24/24) | 20.83% (5/24 flips) | 0.250 | 0.1721 | $0.3244 |
glm-5 |
100.0% (24/24) | 20.83% (5/24 flips) | 0.375 | 0.1581 | $0.6860 |
gpt-5.4-mini-2026-03-17 |
100.0% (24/24) | 25.00% (6/24 flips) | 0.333 | 0.1370 | $0.0591 |
deepseek-r1-0528 |
95.8% (23/24) | 37.50% (9/24 flips) | 0.042 | 0.1967 | $0.4244 |
qwen3-next-80b-a3b-thinking |
100.0% (24/24) | 41.67% (10/24 flips) | 0.208 | 0.1816 | $0.1753 |
Domain Breakdown: Where Models Failed
Disaggregating results across sectors highlights sharp contrasts between model families:
| Model Identifier | Education PVR | Healthcare PVR | Justice PVR | Finance PVR | Healthcare Sycophancy | Education Sycophancy |
|---|---|---|---|---|---|---|
gemini-3-flash-preview |
0.00% | 16.67% | 16.67% | 33.33% | 0.167 | 0.333 |
qwen3-next-80b-a3b-instruct |
0.00% | 0.00% | 16.67% | 50.00% | 0.333 | 0.167 |
claude-sonnet-5-default |
33.33% | 16.67% | 33.33% | 0.00% | 0.167 | 0.333 |
gemini-3.7-flash |
33.33% | 16.67% | 33.33% | 0.00% | 0.167 | 0.333 |
glm-5 |
0.00% | 16.67% | 33.33% | 33.33% | 0.167 | 0.667 |
gpt-5.4-mini-2026-03-17 |
16.67% | 16.67% | 33.33% | 33.33% | 0.500 | 0.333 |
deepseek-r1-0528 |
16.67% | 66.67% | 33.33% | 33.33% | 0.000 | 0.000 |
qwen3-next-80b-a3b-thinking |
33.33% | 50.00% | 66.67% | 16.67% | 0.333 | 0.333 |
Core Revelations
1. The Reasoning Model Bias Paradox
Reasoning models produced the highest decision flip rates on the leaderboard. Qwen Thinking reached 41.67% total flips. DeepSeek-R1 reached 37.50% total flips.
Extended internal thinking tokens acted as rationalization engines. When demographic tokens changed, internal reasoning chains generated extra assumptions regarding flight risk or family poverty. In criminal justice, Qwen Thinking flipped 4 out of 6 rulings. It granted bail to dominant-caste applicants while requiring surety bonds from marginalized applicants.
2. The Thinking Versus Instruct Trade-Off
Comparing Qwen Thinking against Qwen Instruct exposed a fundamental trade-off between deliberation and certainty.
Qwen Instruct produced only 4 demographic flips across 24 disputes, achieving 16.67% parity violations. However, it failed ambiguity detection in 12 out of 24 cases. In financial and judicial cases, it rushed to decide outcomes despite missing documentation.
Qwen Thinking achieved a perfect score on ambiguity detection. It correctly refused to decide when records lacked critical context. Yet its internal reasoning tokens later rationalized opposing verdicts when names changed. Thinking tokens prevented premature conclusions while enabling demographic rationalization.
3. The Persona Sycophancy Split
DeepSeek-R1 scored near-zero sycophancy across all four sectors. It maintained its stance regardless of whether hospital directors or activists asked the questions.
Conversely, GLM-5 exhibited strong sycophancy in education. It achieved zero demographic flips under direct testing. Yet when conversational personas engaged the model, it shifted positions entirely. It agreed with reservation critics, then agreed with quota advocates on identical legal facts.
4. The Finance Divide
Western frontier models showed strong consistency in financial tasks. Claude Sonnet 5 and Gemini 3.7 Flash achieved zero decision flips across all financial cases.
However, justice evaluations triggered high violation rates across five models. Statutory protections in criminal law and affirmative action disputes remain difficult for current frontier models.
What I Would Measure Next
- Regional Language Testing: Translating the 24 disputes into Marathi, Konkani, and Hindi to evaluate tokenizer bias.
- Dynamic Argument Graphs: Clustering model reasoning chains directly inside Kaggle notebooks using sentence embeddings.
- Court Precedent Grounding: Testing whether providing official legal statutes removes demographic decision flips.
My Benchmark
You can inspect our live benchmark leaderboard, evaluate model runs, and explore the underlying code on Kaggle:
- Kaggle Benchmark Leaderboard: https://www.kaggle.com/benchmarks/kanakwaradkar/indocentric-institutional-ai-bias-benchmark/leaderboard
- Kaggle Benchmark Suite: https://www.kaggle.com/benchmarks/kanakwaradkar/indocentric-institutional-ai-bias-benchmark
- Kaggle Benchmark Task: https://www.kaggle.com/benchmarks/tasks/kanakwaradkar/indocentric-deep-bias-benchmark/8
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.

