Does your model understand India?
IndiaSocialBench measures how language models respond to emotionally difficult Indian conversations. It evaluates indirect refusals, family negotiations, honor and shame, grief etiquette, and money between friends in English, Hinglish, and Hindi. Every score links to the transcript and judge explanation behind it.
| Model | Overall ↓ | |
|---|---|---|
| 1 | Claude Fable 5 Anthropic · reasoning: capped low | 9.47 |
| 2 | Claude Opus 5 Anthropic · reasoning: capped low | 9.42 |
| 3 | Kimi K3 Moonshot · reasoning: capped low | 9.34 |
| 4 | GPT-5.6 Sol OpenAI · reasoning: capped low | 8.78 |
| 5 | Qwen3.7 Max Alibaba · reasoning: capped low | 8.74 |
| 6 | Claude Sonnet 5 Anthropic · reasoning: capped low | 8.64 |
| 7 | Grok 4.5 xAI · reasoning: capped low | 8.35 |
| 8 | MiniMax M3partial · 42/50 MiniMax · reasoning: capped low | 8.27 |
| 9 | Gemini 3.5 Flash Google · reasoning: capped low | 7.97 |
| 10 | MiMo V2.5 Pro Xiaomi · reasoning: capped low | 7.84 |
| 11 | GPT-5.6 Luna OpenAI · reasoning: capped low | 7.78 |
| 12 | Gemini 3.1 Flash Lite Google · reasoning: capped low | 7.61 |
| 13 | Inkling Thinking Machines · reasoning: capped low | 7.56 |
| 14 | Nex N2 Propartial · 46/50 Nex AGI · reasoning: capped low | 7.47 |
| 15 | GLM-5.2 Zhipu · reasoning: capped low | 7.31 |
| 16 | DeepSeek V4 Pro DeepSeek · reasoning: capped low | 7.27 |
| 17 | DeepSeek V4 Flash DeepSeek · no hidden reasoning | 7.25 |
| 18 | Hy3partial · 46/50 Tencent · reasoning: capped low | 7.21 |
| 19 | Qwen3.6 Flash Alibaba · reasoning: capped low | 7.07 |
| 20 | Sarvam 105B (high reasoning) Sarvam AI · reasoning: high | 7.04 |
| 21 | Mistral Large 2512 Mistral · no hidden reasoning | 6.80 |
| 22 | Claude Haiku 4.5 Anthropic · reasoning: capped low | 6.72 |
| 23 | Grok 4.3 xAI · reasoning: capped low | 6.50 |
| 24 | Sarvam 30Bpartial · 42/50 Sarvam AI · reasoning: capped low | 5.89 |
| 25 | Sarvam 30B (high reasoning) Sarvam AI · reasoning: high | 5.86 |
| 26 | Sarvam 105B Sarvam AI · reasoning: capped low | 5.54 |
| 27 | GPT-5 Mini OpenAI · no hidden reasoning | 5.48 |
| 28 | Llama 4 Maverick Meta · no hidden reasoning | 3.23 |
Scores range from 0 to 10. The whisker on the Overall bar shows the bootstrap 95% CI. “Gap en→hi” is how much the model loses when the same conversations arrive in Hindi (−) or gains (+). Refusals are excluded from scores and reported separately. Click any row for per-dimension detail and full transcripts.
Reasoning policy: to keep runs comparable and affordable, reasoning-capable models are run with thinking effort capped at “low” (marked per model above); rows labeled “high reasoning” are explicit higher-effort variants, shown separately so no model gets a hidden advantage.
What the numbers say
The pilot shows a language gap
18 of 28 models score lower when the identical situations arrive in Hindi instead of English. The average drop is 0.32 points. Same problems, same rubric, same judges; only the language changed.
Culture is harder than empathy
The field averages 8.55 on the culture-neutral support control, but 7.20 across the seven cultural dimensions and is weakest on indirect speech & face-saving (5.96). Models know how to feel; they don't yet know how India works.
Refusals on family topics
MiniMax M3 refused ordinary family-life scenarios at 4% or more. Over-refusal is reported as its own column and is never hidden in the average.
Excluded from this board: GLM-4.7 (only 19/50 items completed, insufficient coverage (provider errors)).
Dataset vd7bb3900ca057116 · generated 2026-07-26 · 18 base scenarios · single uniform blinded judge (see methodology & limits)