← All projects

Violet

Smart glasses that help people with dementia recognize the people they love.

2026
  • Meta Glasses SDK
  • AWS Rekognition
  • MongoDB Atlas
  • Python
Open demo ↗

Inspiration

Around 57 million people live with dementia, with nearly 10 million new cases each year (World Health Organization, 2026). There is no cure, but early intervention and continued support — cognitive stimulation, social engagement, medication — can help improve daily functioning.

One of the most painful parts is losing the thread with the people you love. Violet doesn’t try to remember for someone. It hands them context — a relationship, a recent visit, a shared story — so they can reason their own way back to the memory. The cue is itself cognitive stimulation, and it keeps them engaged with the person in front of them.

It supports the people around them too: with the patient’s consent, a caregiver portal turns everyday usage into trends a clinician can read.

What it does

Violet pairs Meta smart glasses with an iPhone app that recognizes familiar people and tells the wearer who they are, plus a portal where caregivers and clinicians follow how the patient is doing.

A caregiver first enrolls familiar people, with their consent: three reference photos each, plus a name, their relationship to the patient, and any notes or memories worth surfacing.

When the wearer says “Hey Violet,” the glasses capture whoever is in view. Violet matches the face, pulls up that person’s details, and a language model condenses them into a short cue that deliberately leaves gaps for the wearer to fill. Instead of “This is Mark, your grandson, who started a new job as a radiologist at Emory Hospital six days ago,” Violet says “This is Mark, your grandson, who told you about his job last week.” A follow-up question can ride along (“Hey Violet, where does he work?”) and gets its own answer after the cue.

Forgetting someone can be embarrassing, so there is a silent path too: press the capture button on the glasses, or tap a face in the app to hear that person’s bio.

How we built it

The recognition pipeline

In a live conversation, latency, cost, and battery all matter. The glasses don’t stream constantly: Meta’s on-device speech recognition listens for the wake phrase, and only then does Violet start pulling video.

Sending every frame to the cloud would be slow and expensive, so the phone does the first pass. Frames arrive at 15 FPS; Apple Vision finds faces and landmarks, and a small on-device quality model scores each crop and throws out the weak ones. Only the strongest crop per face goes to AWS Rekognition, which returns the best-matching enrolled identity with a confidence — or no match at all.

When several enrolled people are in view, a deterministic score picks the one the wearer meant, weighing how centered and how large each face is and how close to the wake phrase it was seen.

Recognition runs while the glasses are still capturing. Each new face’s best crop goes out immediately, with several Rekognition calls in flight at once. As soon as Violet is confident it ends the capture early, and a hard five-second deadline means a slow network never leaves the wearer waiting.

Choosing the on-device model

The quality model exists to cut unnecessary Rekognition calls without paying the latency and battery cost of something heavy. We fine-tuned pretrained face-recognition backbones and attached a small MLP head that takes the pre-normalization face embedding and the crop’s effective resolution, and predicts how useful that crop will be to Rekognition.

Three backbones of increasing size mapped out the accuracy-versus-efficiency tradeoff:

  • MobileFaceNet — ~2M parameters
  • EdgeFace-S — ~3.65M parameters
  • ArcFace iResNet-50 — ~44M parameters

Each ran both frozen and partially fine-tuned, across a 162-configuration hyperparameter sweep on two T4 GPUs.

Results table. MobileFaceNet fine-tuned: rank correlation 0.63, AUROC 0.79. EdgeFace-S fine-tuned: 0.60, 0.76. iResNet-50 fine-tuned: 0.60, 0.75. Best frozen model: 0.49, 0.74.

Every fine-tuned variant beat its frozen counterpart, and the smallest, MobileFaceNet, scored best of all. The most accurate model is also the cheapest to run, so the quality check runs on the phone in under a second.

Building the dataset

There is no off-the-shelf label for “how well will Rekognition do on this crop,” so we made one. We took 570 CelebA identities, split by person into train, validation, and test sets, enrolled three reference photos of each into Rekognition, and generated about 10,000 query crops, roughly half synthetically degraded to simulate blur, poor lighting, compression, and sensor noise. Every query went through Rekognition, and its actual recognition performance became the training signal.

Voice and audio

The glasses transcribe speech on their own, so the wearer only ever has to say “Hey Violet.” A short chime confirms they were heard, and the answer — voiced by ElevenLabs — plays through whatever output is active, including the glasses’ speakers. There are four possible answers: a named person, someone who isn’t enrolled, no face in view, and an unsure match, where Violet declines to guess.

Waiting on text-to-speech after recognition would add a noticeable pause, so every sentence Violet could say is generated and cached ahead of time; when someone’s bio changes in the portal, their lines regenerate in the background. If an answer still takes longer than three seconds, Violet fills the silence with a short “One moment.”

The provider portal

The portal is for caregivers and clinicians. It’s built with Next.js on MongoDB Atlas and synced with the patient’s Google Calendar, so recognitions can be matched to scheduled visits.

It shares one Atlas database with the patient’s phone, so a person added or edited in the portal reaches the glasses within a minute; the phone then regenerates that person’s voice lines and re-enrolls their photos in Rekognition. In the other direction, every recognition is logged on the phone, uploaded to Atlas, and matched against calendar events that name an enrolled person. Both sides keep a local cache, so the phone works offline and the portal renders instantly before refreshing in the background.

Clinicians can also keep rich-text notes on the patient, alongside their caregiver contact and latest MoCA score.

Designing for people with dementia

iOS app: how do we keep the interaction simple for the patient?

  • There is one thing to remember: say “Violet,” or press the button on the glasses.
  • Violet never guesses. It names someone only on a high-confidence match, since a confident wrong answer would confuse someone who can’t double-check it.
  • Answers carry context, not just a name: “This is Jordan Lee, your daughter.”
  • There is deliberately no way to delete a person from the phone, so a confused patient can’t remove family members.

Provider portal: what information is actually useful to a caregiver or doctor?

  • Recognition health — at each scheduled visit, did Violet identify the person on the calendar? A mismatch can come from a wrong calendar or an unrecognized face, so this tells caregivers whether the rest of the data can be trusted.
  • Weekly trend — uses per visit over time, a read on how the condition is progressing.
  • Time of day — calls clustering late in the day can signal sundowning (late-day confusion and agitation), so caregivers can anticipate high-risk windows for wandering and plan extra support.
  • Memory by person — calls plotted against how long the patient has known each person. Dementia usually takes recent memories first, so frequent calls about newer acquaintances fit the expected pattern; when long-known family start coming up too, the disease may have progressed.

System architecture

Two clients share one database: the Swift app on the patient’s phone and the Next.js provider portal both sync through Atlas, with Rekognition, ElevenLabs, and Google Calendar at the edges. The quality model is trained offline in Python and exported to Core ML to ship inside the app.

Violet system architecture: Meta glasses and the patient iOS app, the Rekognition, Atlas, ElevenLabs and Google Calendar services, the Next.js provider portal, and the offline ML pipeline that exports the Core ML face-quality model

Challenges we ran into

  • Building a useful dataset. Training the quality model needed clean reference images with landmarks and hard queries — bad angles, blur, compression, poor light. Datasets that look like that (SCFace, for one) sit behind access requests we couldn’t clear during a hackathon, so we used CelebA and synthetically degraded about half the queries to approximate a wearable camera. Even then, Rekognition’s scores bunched up near 100, so a logarithmic transform stretched out the top end to make the differences learnable.
  • Designing for someone who may forget new interfaces. Most apps assume people pick up features over time; our users might forget anything advanced on day one. That’s why the whole interaction comes down to saying “Violet” or pressing one button.
  • Keeping pace with a real conversation. Even a few seconds of silence is awkward with someone standing in front of you, and the glasses camera alone takes about a second to deliver its first frame. So everything that can overlap does: faces are scored on-device as frames arrive, Rekognition calls start mid-capture and run in parallel with an early exit on a confident match, and all speech is generated ahead of time.

What we learned

  • Bigger isn’t better by default. Pick model complexity for the data and the task. On a limited dataset, the smallest backbone generalized best and was also the easiest to ship on-device.
  • Start from the user’s limits. The good product calls came from the needs of people with dementia: simple interfaces, no unnecessary features, familiar interactions like voice, even a human name like “Violet” so the assistant feels approachable.
  • Several narrow models beat one big one. Face-quality scoring, identity matching, referent selection, and conversational context each solve a narrower problem more reliably than one giant model asked to do everything, and keeping them separate made each easier to test, debug, and build in parallel.

What’s next for Violet

  • Automatic memory updates — with consent, briefly transcribe the conversation after Violet recognizes someone and let a language model pull out details worth adding to their notes, so the next cue is fresher.
  • Adaptive cues — learn how much a particular user needs before a name clicks, and add detail only when it’s needed instead of giving the same-size answer every time.
  • Multimodal follow-ups — let the model answering follow-up questions reason over what the glasses see, not just the notes.