Teaching a Wyze Cam v3 to talk like a Skyrim guard
The plan: put open firmware on a $35 Wyze camera, run everything on local AI in my homelab, and give it personalities, from a Skyrim guard to a very nervous anime character.
I have a Wyze Cam v3 and a homelab with a GPU in it. So naturally, the camera is going to become a character.
The goal: when someone walks past, the camera notices, says something in character through its speaker, and maybe reacts to what they’re doing. Personas I want to try:
- Skyrim guard. “Let me guess, someone stole your sweetroll?”
- Nervous anime character who is very sorry for noticing you.
Everything runs locally. No cloud, no Wyze servers, no recordings leaving my house. That’s half the point.
The plan
1. Free the camera. The stock firmware sends everything to Wyze’s cloud. Open-source projects like Thingino replace it with firmware that gives you a local RTSP video stream and full control. That’s step one.
2. Watch the stream locally. A service on my server pulls the stream and runs lightweight motion or person detection, so the expensive AI only runs when something is actually happening.
3. Describe and react. When someone shows up, a local vision model describes the scene, and a local language model writes a line in the current character’s voice. All of it runs on my own hardware.
4. Give it a voice. A local text-to-speech model turns the line into audio. Then I get it playing out of the camera’s own speaker, which is the part I’m least sure about.
Unknowns
- How much audio control the open firmware gives me on the v3
- How fast the whole loop is. A guard who takes 15 seconds to respond is less intimidating.
- Whether my family will tolerate this
I’ll post each step here and on Instagram as it comes together, including the parts that don’t work.
Follow the project on Instagram (@tarchosit). Need something like this fixed? Here's how I can help.