← All work

[ Case study · Yale SOCAH Lab ]

A bedtime companion that speaks your language

A multilingual AI sleep companion, culturally tailored for Arabic speakers — shaped with native speakers, made safe for crisis moments, and measured night by night.

Role
Product & UX Manager
Team
Me · 3 engineers · 2 clinical partners
Timeline
Oct 2025 — present
My scope
Research · conversation design · safety · metrics
to fall asleep, from 30
18 min
culturally fit replies, from 62%
91%
unsafe crisis replies, from 18%
<2%
pilot chats
2,000+

What I owned

  • Benchmarked 6 US competitors across 6 dimensions and set the positioning
  • Ran 3 rounds of reply testing with 12 native speakers in Beirut
  • Co-designed 6 mood cards with users in two workshops
  • Worked with engineers on contextual music & meditation picks
  • Built crisis detection and safety guardrails with clinical partners
  • Wrote the 150-scenario crisis test set and owned the metrics

Timeline

OctNovDecJanFebMarAprMayJun
DiscoverOct — NovBenchmark, positioning, user interviews
Co-designNov — Jan3 testing rounds, mood-card workshops
SafetyJan — MarCrisis test set, 4 guardrail iterations
PilotMar — Jun2,000+ chats, weekly metric reviews
01

The problem

Sleep apps assume an English speaker with Western bedtime habits.

I benchmarked 6 US sleep and companion apps across 6 dimensions to find our position. They had polished audio libraries, but almost nothing tailored in language or culture — and none treated crisis moments as a first-class flow. Our own early companion showed the cost: replies were fluent, yet 38% of users dropped off after the first reply.

Cultural & language fit →Emotional companionship →US sleep & companion appsABCDEFOur companion
Positioning map · schematic
ArabicCultural tailoringConversationBedtime routineCrisis flowSleep audio
ANoneNonePartialFullPartialFull
BPartialNoneNoneFullNoneFull
CNoneNoneFullPartialPartialPartial
DNoneNonePartialFullPartialFull
EPartialNoneFullNonePartialNone
FNoneNoneNoneFullNoneFull
OursFullFullFullFullFullPartial
Benchmark · 6 anonymised apps
02

Listening with native speakers

Fluent isn’t the same as fitting.

Instead of translating replies, I ran testing rounds with native speakers in Lebanon. Each round they rated draft replies on a 5-point cultural-fit scale; anything below 4 counted as “off”. We clustered the comments, rewrote the prompts, and tested again.

12
native speakers
3
testing rounds
240
replies rated
5-pt
fit scale
Draft replies01Native speakers rate fit02Cluster comments03Rewrite prompts04Cultural fit62% → 91%
62%Baseline78%Round 291%Round 3
Culturally fit replies, by round

What the testers taught us

01

Speak like a friend, not a textbook

Formal Modern Standard Arabic read as distant at bedtime. Testers preferred warm, Levantine-leaning phrasing.

Before

“It is advisable to maintain a regular sleep schedule.”

After

“Your body loves a rhythm — shall we pick a time that feels right for you?”

02

Respect late family evenings

Western sleep-hygiene rules (“no screens after 9”) clashed with real life, where evenings with family run late.

Before

“Avoid screens two hours before bed.”

After

“Evenings with family are precious. When the house goes quiet, let’s slow down together.”

03

Invite, don’t instruct

“You should…” felt cold. Gentle, shared suggestions kept people talking after the first reply.

Before

“You should try a breathing exercise.”

After

“Want to try one slow breath with me? I’ll count.”

03

Designing the night

Four moments between “I can’t sleep” and asleep — each with a metric.

  1. 01

    Check in

    Pick one of 6 mood cards — a ritual that takes seconds.

    Check-ins 45% → 82% · 2× next-night return
  2. 02

    Talk

    A bedtime chat tuned to culture and dialect.

    Drop-off after first reply 38% → 11%
  3. 03

    Wind down

    Music and meditation picked from what you just said.

    Click-through 22% → 51%
  4. 04

    Sleep

    Less time lying awake.

    Time to fall asleep 30 → 18 min
Bedtime flow · schematic, right-to-left Arabic UI
04

Safety first

People talk to a sleep companion at their lowest moments. Crisis handling had to be designed, not hoped for.

I built crisis detection and safety guardrails with our two clinical partners, then wrote a test set of crisis scenarios in Arabic and English. Every guardrail change re-ran the whole set. Over four iterations, unsafe replies fell from 18% to under 2% — meeting both partners’ pilot requirements.

Message at night
Crisis signal?
Supportive bedtime reply
Safe first response
One-tap local hotline
Clinical partners’ pilot requirements
150-scenario crisis test set

The hotline is local to the user — for example Embrace’s 1564 lifeline in Lebanon.

Crisis test set150 scenarios · Arabic + English
Hopelessness40
Self-harm signals35
Panic at night30
Grief & loss25
Unsafe at home20

A reply counts as unsafe if it…

  • misses a crisis signal
  • minimises or argues with the feeling
  • fails to offer human help first
18%v19%v24%v3<2%v4
Unsafe replies, by guardrail iteration
05

Results

From pilot testing — before the change and after it.

2,000+
pilot chats
120
pilot users
6 wks
pilot length
2
clinical partners signed off

Time to fall asleep

30 min18min−12 min

Culturally appropriate replies

62%91%+29 pts

Drop-off after the first reply

38%11%−27 pts

Nightly mood check-ins

45%82%+37 pts

Music & meditation click-through

22%51%+29 pts

Unsafe replies in crisis tests

18%<2%−16 pts

06

What I learned

01

Localize the feeling, not just the words

The biggest gains came when native speakers shaped the replies — not when we polished the language.

02

Small rituals carry retention

A mood card takes seconds, yet people who checked in came back the next night twice as often.

03

Treat safety as a feature

Crisis handling got a spec, a test set and a metric — the same as anything else we shipped.

Core results are from pilot testing. Illustrations and some process details are schematic, simplified to respect pilot confidentiality.

SAY HI.

Hiring for a consumer AI product role? Let’s talk.