Agentic AI nighttime caregiver
I took overnight dementia care from need-finding to a working, evaluated prototype. Here's what I built and what I learned.
- Role
- Solo builder
- Context
- Stanford Biodesign
- Years
- 2025–2026
- Status
- Completed prototype
Problem
About one in five people with Alzheimer's who live at home wake up frequently at night. In the US alone, that's roughly 800,000 people. When they wake up confused or agitated, the person caring for them wakes up too, sometimes several times a night. This is one of the hardest symptoms for the family caregivers since this turns daytime caregiving to full 24h shifts without breaks, often leading to caregiver burnout .
Families mostly have two options today. Some put smart-home cameras in the bedroom and living room, but someone still has to watch the feed or get up for every alert. Others move their relative into senior living, which averages around $6,200 a month in the US (CareScout 2025 Cost of Care Survey).
Additionally, we learned that >80% of the reason patients wake up at night are non-physical in nature. Patients tend to wake up, start talking, wander in their bedroom or get lost on the way to the bathroom. Less than 20% is diaper changes or falls, which are things that absolutely require a physical intervention.
Through 50+ interviews with caregivers, facilities and clinicians, we found that families would pay over $200 a month out of pocket for something that reliably helped reduce caregiver nighttime awakenings.
My role
The need-finding and validation happened during my Stanford Biodesign Innovation Fellowship, as part of a team. Together we ran needs finding, screening and concept selection, and worked out a regulatory and commercialization plan on how to reduce reducing caregiver night-time awakenings in Alzheimer's.
After the fellowship I built the system on my own as the only engineer. That covered perception, the dialogue agent, the caregiver dashboard and the evaluation. Healthy volunteers contributed RGB camera recordings of themselves acting out different night-time scenarios, which I used to test and tune detection.
What I did
The system watches a bedroom, detects distress and falls, speaks reassurance through a bedside display, and alerts caregivers in real time. Perception runs fully on-device (under 100 ms per frame on an M4 Mac, sampled at 2 fps), so bedroom video and audio stay off the cloud. Services talk over a Redis Streams bus, which also let me record and replay whole nights for hardware-free testing.
Interestingly, we can also annotate the dataset using a VLM model (Vision-Language Models (VLMs)). Here, annotation takes multiple seconds per frame, so this cannot be used for real-time monitoring but could be used for a privacy-focused analysis of frames/nightime activities.
Remote access is TLS-secured, and a caregiver dashboard handles zone configuration, event history and profile management.
The design goal is an embodied agent that guides patients back to sleep with a familiar-voice, low-arousal audiovisual cues, adapts to their preferences over time, and escalates to a caregiver only when someone needs to step in.
Result - Detection
I tested detection on 5+ scripted bedroom recordings of healthy volunteers, with 46 annotated events in total, 6 of them falls to the floor. Tuning the pose model and adding simple geometric rules for where the floor is brought on-floor recall to 94–100%, and floor events were flagged within the 2-second alert target. I replayed the same recordings after every change, so each result could be compared directly with the last.
| Measure | Result |
|---|---|
| Annotated events | 46 |
| Falls to the floor | 6 |
| On-floor recall, after tuning | 94–100% |
| Alert latency target | 2 s, met for floor events |
| Perception speed (M4 Mac, 2 fps) | < 100 ms per frame |
Scripted testing with healthy volunteers. Not a clinical validation.
Result - Agentic AI
Nobody can test a night companion at 3 a.m. with a real person in the room, so I built simulated nights. Claude Opus writes each scene, and when a scene breaks something it writes a smaller variant to isolate the cause. A second Claude agent plays the person through real synthesised speech, and the real speech recognition, agent and bedside voice answer in real time. Fifteen automatic checks read every night's trace: did each question get a reply within 5 seconds, did the agent talk over the person, did it fall silent when it shouldn't.
Opus then triages the night. It traces each cluster of failures to the code, tries a fix in a throwaway copy, replays the failing moment and writes a ranked fix list with questions for me. Nothing is applied until I say so. In one simulated fall, a woman lying on the floor was told "take your time getting back to bed", twice. The fix was a new rule that never sends a return-to-bed prompt to someone on the floor. Before it, that prompt went out 3 times in 8 floor scenes; across the four runs since, it went out 0 times in 18.
| Measure | Result |
|---|---|
| Simulated hours with automatic triage | 5 |
| Simulated scenes | 74 |
| Flags raised | 387 |
| Ranked fixes proposed | 37 |
Simulated people, real agent stack and local model. The rules rest on clinical guidance I checked source by source; this is not a clinical validation.