Local proxy that gives vision to text-only models (DeepSeek V4 Flash and any other). A vision model describes your screenshots — text, hex colors, layout — and forwards them as text to your main model. Works with Trae, Cursor, OpenCode, Codex, Claude Code and any OpenAI/Anthropic-compatible client. Dockerized.