Local video search with Qwen3-VL: no API, runs on Apple Silicon, GPUs
Local video search with Qwen3-VL: no API, runs on Apple Silicon, GPUs
Last week, I posted SentrySearch, a CLI for semantic video search using Gemini's embedding API. The #1 request was local model support. Turns out Qwen3-VL-Embedding can natively embed video into the same kind of vector space, no API, fully offline. Runs on Apple Silicon (MPS) and NVIDIA GPUs (CUDA). The 8B model needs ~18GB RAM, or use the 2B model on smaller machines. sentrysearch index /path --backend local Also added: similarity threshold to suppress weak matches, and a Tesla metadata overlay that renders speed/location onto matched clips. Details on the README.
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Correct prediction on native model
Similar products
ArrowMetal – Apache Arrow Compute on Apple Silicon GPUs via Metal
I made Android boot on Apple Silicon
MLX (Apple Silicon tensor library) bindings for Erlang
NLP PyTorch Tutorial (fire up the GPUs)
Heroku for GPUs
Fractional GPUs for AI
Silicon Feelings
local speech-to-text is shockingly fast on Apple Silicon
Shadeform – Single Platform and API for Provisioning GPUs
Oceananigans.jl: A fast ocean model in Julia that runs on CPUs and GPUs