I made both of them watch 165 minutes of an Indian comedy quiz show, loaded with Indian pop culture references, Hinglish speech and dialogue that only desi people would understand. Here's who understood India better, and at what cost.
In the show, Gursimran Khamba, the host, asks three comedians weird questions about real news and Indian pop culture. It's all in Hinglish, at the speed of banter. It can't get more desi than this.
An AI that can follow this is definitely more desi.
the hostthe guessersthe audience
Gursimran Khamba host
Abish Mathew, Rohan Joshi, Kaneez Surka guessers
01
Job 1 of 6
Play the game, as contestants.
Gursimran's own questions, every rule followed. This is where "more desi" gets decided.
Video 1 · 58:42
Round 3 · Who am I?
I was virtue signalling before it was cool
Haters prayed for my downfall but I kept slaying
Every Miss World would name drop me in the 90s
Jose buzzesSunil Pal✗
Sarvamlistening…
OpenAIlistening…
Both AIs played the same hints. Hint 3 is where they split: that's Round 1.
02
Job 2 of 6
Then make sense of the game itself.
The rounds, the rules, the scores, the wagers, and who actually won.
True Story
Audience round
Pick Me Behaviour
Who Am I?
0Abish
0Rohan
0Kaneez
The real Video 2 scoreboard: after Round 1, before the last round, then the all-in wager
03
Job 3 of 6
Watch both episodes, scene by scene.
165 minutes of video: who's on stage, what the cards say, what's on the screen.
Video 1 · 75 minVideo 2 · 90 min
04
Job 4 of 6
Write down every word.
In Hinglish, exactly the way it was said. Two languages, one sentence, at the speed of banter.
Video 2 · 77:02
Gursimran KhambaI am the aakhri pasta. I am the?
A guesserAakhri pasta. Aakhri. Aakhri, last.
A guesserThe last pasta.
05
Job 5 of 6
Understand the audience's reaction.
On this show the claps are points. The AI has to hear every laugh and every applause break, because that's how the game is scored.
+20
Video 2 · 2:54 · the hostWe are India's most unserious comedy quiz show, jahan right answer ko milta hai plus ten, and applause break ko milta hai plus twenty. Show me the applause break!
+10 right answer+20 applause break
06
Job 6 of 6
Figure out who's talking, by name.
The host and three comedians, talking over each other. Every line needs the right name on it.
Video 2 · 77:17
Gursimran
Abish
Rohan
Kaneez
Names only where both AIs agree
Question 1
Who understands Indian pop culture better?
Real moment · Video 1 · 39:20 · the show's own question“In front of which bakery in Bandra did Salman Khan's Land Cruiser have its accident?”
Sarvam
“A1 Bakery”✗
000
+20OpenAI takes it
OpenAI
“American Express Bakery”✓
020
Question 2
Who can understand the video better?
Real moment · Video 1 · 43:50On screen: three comedians at their podiums.
Sarvam
“The image shows two men standing behind podiums…”✗
000
+20OpenAI takes it
OpenAI
“Jose Covaco, Kenny Sebastian, Abish Mathew”✓
040
Question 3
Who hears Hinglish better?
Real moment · Video 2 · 65:03A guest says: “Here's a suggestion: get a job.”
Sarvam
“Here's a suggestion get a job.”✓
020
+20Sarvam takes it
OpenAI
wrote nothing✗
040
Question 4
Who hears the audience laugh?
Real moment · Video 2 · 50:50Morarji Desai is revealed and the whole room erupts.
Sarvam
heard nothing✗
020
+20OpenAI takes it
OpenAI
“LAUGHTER, 3 of 3”✓
060
Question 5
Who knows who's talking?
Real moment · Video 2 · all 90 minutesOne host, Gursimran Khamba. How many people did each AI think he was?
Sarvam
kept him as 2 voices, all show long✓
040
+20Sarvam takes it
OpenAI
split him into 12 voices✗
060
Question 6
Who's cheaper to run?
Real moment · every API call, loggedOne hour of the show, all six jobs. What does it cost?
Sarvam
$5.50 an hour3×
040
+20OpenAI takes it
OpenAI
$1.84 an hour✓
080
1
Round 1 of 6
Who understands Indian pop culture better?
Slang, references, jokes, and Gursimran's own quiz. The "more desi" question.
Sarvam
Sarvam 105B
Sarvam's own LLM
VS
OpenAI
GPT-6 Luna*
OpenAI's LLM
*GPT-6 Luna: I picked the GPT-6 variant priced like Sarvam 105B (both under $1 per million tokens). Putting OpenAI's top model against it would be an unfair fight.
Both get only the hints, exactly like the blindfolded comedians. This is the question in the title.
Who Am I? The show's toughest round. Gursimran reads the hints; both AIs play along.
Video 1 · 58:42both AIs are thinking…
0:00
Who am I?
Hint 1I was virtue signalling before it was cool
Hint 2Haters prayed for my downfall but I kept slaying
Hint 3Every Miss World would name drop me in the 90s
AnswerMother Teresa
Who am I?
Hint 1They see me rolling, they hating
Hint 2Battery life is very important
AnswerStephen Hawking
Who am I?
Hint 1Everyone asks me, "mummy kaisi hain"
Hint 2Haq se single
Hint 3Winning isn't everything
AnswerRahul Gandhi
*GPT-6 Luna: the GPT-6 variant priced like Sarvam 105B. Both AIs only get the hints, like the blindfolded comedians.
Round 1 · resultWho understands Indian pop culture better?
OpenAIwins
MORE DESI: OPENAI
First blood, and the question in the title
Sarvam
0Who Am I points, of 270
vs
Winner+20
OpenAI
90Who Am I points, of 270
higher is better
So, which AI is more desi? On this show, OpenAI. It also knew more slang (94% vs 59%) and more of Gursimran's quiz (10 vs 6.5 of 15).
0after 1 of 620
2
Round 2 of 6
Who can understand the video better?
How many people are on screen, what the cards say, and does it make things up.
Sarvam
Sarvam Vision
⚠ LimitationSarvam Vision is a document reader. It can't be asked a question about a scene, so all it returns is the text it sees or a general caption.
VS
OpenAI
GPT-6 Luna*
looks at the frame and answers questions about it
*GPT-6 Luna: I picked the GPT-6 variant priced like Sarvam 105B (both under $1 per million tokens). Putting OpenAI's top model against it would be an unfair fight.
Five scenes from the show. What each AI said it saw, in the exact frame it was shown.
Video 1 · 43:521 / 5
0:00
On screenThree comedians at their podiums: Guesser 1, 2 and 3.
*GPT-6 Luna: I picked the GPT-6 variant priced like Sarvam 105B (both under $1 per million tokens). Putting OpenAI's top model against it would be an unfair fight.
60 random frames, each one checked
Sarvam reads. OpenAI sees.
Counted the people right
51%89%
Described the scene right
60%90%
Read the on-screen text right
82%85%
SarvamOpenAI
Made things up
7OpenAI
0Sarvam
Only one of them never invented anything. OpenAI also names who's on stage (81% right); Sarvam can't.
Round 2 · resultWho can understand the video better?
OpenAIwins
OpenAI doubles up
Sarvam
51%people counted right
vs
Winner+20
OpenAI
89%people counted right
higher is better
Sarvam reads. OpenAI sees.
0after 2 of 640
3
Round 3 of 6
Who hears Hinglish better?
Who writes down more of what was actually said. I count every word it gets wrong.
Sarvam
Saaras v4
speech to text · tuned for Hinglish
VS
OpenAI
GPT Transcribe
speech to text
Five real moments, with sound. What was said, and what each AI wrote, word for word.
Video 2 · 77:051 / 5
0:00
What was saidGursimran Khamba: I am the aakhri pasta. I am the?
Not cherry-picked
24 clips. 3,130 words. Every one checked.
The same clips through both. For every 100 words actually said, here is every mistake each one made.
Sarvam
Saaras v4 · tuned for Hinglish
18.7%of words wrong
8 wrong words
2 missing words
8 extra words
OpenAI
GPT Transcribe
21.1%of words wrong
5 wrong words
15 missing words
1 extra words
Every 100 squares is 100 words actually said. OpenAI's mistakes are mostly missing words: it skips.Sarvam's are wrong or extra words: it hears everything, not always right.
Round 3 · resultWho hears Hinglish better?
Sarvamwins
Sarvam hits back
Winner+20
Sarvam
18.7%words wrong
vs
OpenAI
21.1%words wrong
lower is better
Sarvam hears more. OpenAI hears cleaner, when it hears at all.
20after 3 of 640
4
Round 4 of 6
Who hears the audience laugh?
Comedy is half the laugh. Can it tell when the room cracks up.
Sarvam
✗No model
⚠ LimitationSarvam has speech-to-text, but no model that listens to sound itself. It can only guess a laugh from the words.
VS
OpenAI
GPT Audio Mini
listens to the audio itself
Three times the room reacted. What each AI heard.
Video 2 · 50:421 / 3
0:00
What was saidKhamba: Abish has picked mera number one. Morarji Desai, apart from being a former PM, was also known for enjoying an occasional drink of his bottle number one.
70laughs and applause breaks heard by OpenAI, across both videos
vs
1by Sarvam, guessed from the words, not heard
Every laugh and applause break OpenAI heard in Video 2
0:0045:0790:15
Sarvam, same 90 minutes
Round 4 · resultWho hears the audience laugh?
OpenAIwins
OpenAI pulls ahead
Sarvam
1laughs heard
vs
Winner+20
OpenAI
70laughs heard
higher is better
Comedy is half the laugh. OpenAI heard the audience. Sarvam couldn't.
20after 4 of 660
5
Round 5 of 6
Who knows who's talking?
Does it keep the host and the three comedians straight, all show long.
Sarvam
Saaras v4
speaker labels, one pass over the whole show
VS
OpenAI
GPT-4o Transcribe Diarize
speaker labels, ten minutes at a time
Three stretches of crosstalk. Every line, with the name each AI put on it.
Video 2 · 63:261 / 3
0:00
Who actually said itGursimran: Yes, he was almost right. Let's give you the answer. BIA. No, no, no, no, no. No, you have a good thing going into it. IB, IB, IB. / A guest: IB, IB, irritable bowel syndrome. / Kaneez: AI, AI. / Gursimran: Yes, Kaneez got the right answer, yes. What? But we should give Abish also a little bit, he helped. Let's give Kaneez ten and Abish five.
All of Video 2, 90 minutes
Four people talking. How many voices did each AI hear?
4people actually talking
8voices, as Sarvam heard it
61voices, as OpenAI heard it: it forgets everyone every ten minutes
Round 5 · resultWho knows who's talking?
Sarvamwins
Sarvam closes in
Winner+20
Sarvam
64%lines on the right person
vs
OpenAI
41%lines on the right person
higher is better
Sarvam kept the voices straight all show long. OpenAI forgot everyone every ten minutes. Forty to sixty, and it comes down to the bill.
40after 5 of 660
6
Round 6 of 6
Who's cheaper to run?
Every rupee, part by part. Listening, clean-up, watching, thinking.
Sarvam
The whole Sarvam stack
Saaras v4 · Sarvam Vision · Sarvam 105B
VS
OpenAI
The whole OpenAI stack
GPT Transcribe · GPT Audio Mini · GPT-6 Luna*
*GPT-6 Luna: I picked the GPT-6 variant priced like Sarvam 105B (both under $1 per million tokens). Putting OpenAI's top model against it would be an unfair fight.
Every API call was logged with what it cost. Two videos, 2.75 hours.
The bill, per hour of video
Two receipts. Same job.
Sarvam
per hour of video · measured
Listeningevery word, and who is talking$0.54
Clean-upfixing the transcript (and, for OpenAI, hearing the room)$0.28
Watching78% of the billreading the screen, frame by frame$4.29
Thinkingmaking sense of the show, minute by minute$0.39
Total$5.50₹462
3× the price
OpenAI
per hour of video · measured
Listeningevery word, and who is talking$1.34
Clean-upfixing the transcript (and, for OpenAI, hearing the room)$0.39
Watchingreading the screen, frame by frame$0.078
Thinkingmaking sense of the show, minute by minute$0.029
Total$1.84₹154
Cheapest
Green amounts: the cheaper of the two, part by part. Sarvam is cheaper at listening; OpenAI at everything else.
Round 6 · resultWho's cheaper to run?
OpenAIwins
The receipt
Sarvam
$5.50per hour of video
vs
Winner+20
OpenAI
$1.84per hour of video
lower is better
Overall, Sarvam cost three times more. Measured on every call, not estimated.
40after 6 of 680
The final score · six tests
OpenAIwinstheshow
Sarvam
40points
vs
Winner
OpenAI
80points
Q1Pop cultureOpenAI +20
Q2The videoOpenAI +20
Q3Hinglish earsSarvam +20
Q4The laughsOpenAI +20
Q5Who's talkingSarvam +20
Q6The billOpenAI +20
Sarvam hears Indian speech better, and cheaper.
OpenAI understands it better, sees better, and costs a third as much overall.