Modular async pipeline: FrameSource->Vision->context and STT/text->Brain->TTS. All stages are Protocols; mock backends run end-to-end with no deps/keys. Real backends included: mss screen capture, Claude vision+brain (guarded imports).
18 lines
435 B
Plaintext
18 lines
435 B
Plaintext
# Core skeleton has NO required third-party deps (mock mode is pure stdlib).
|
|
# Install extras per backend you enable:
|
|
|
|
# --- screen capture (WSAI_SOURCE=mss) ---
|
|
# mss
|
|
# pillow
|
|
|
|
# --- cloud eyes + brain (WSAI_VISION=claude / WSAI_BRAIN=claude) ---
|
|
# anthropic
|
|
|
|
# --- planned voice backends (not yet implemented) ---
|
|
# faster-whisper # STT
|
|
# sounddevice # mic capture
|
|
# (a TTS engine, e.g. MeloTTS)
|
|
|
|
# --- dev ---
|
|
# pytest
|