Voxae
Language-grounded, metric scene understanding for physical AI.
Ask an aerial scene a question in plain language and get the region that answers it. Two systems run on the same query:
- Trained
<SEG>bridge: a vision-language model whose hidden state at a<SEG>token is projected directly into SAM 2.1's prompt space and trained end to end on Voxae-Reason (referring + affordance queries over drone imagery). - Zero-shot baseline: a hosted VLM emits a box and points as text, SAM 2.1 turns them into a mask. No training.
The gap is widest on implicit queries: naming an object is easy, reasoning about what a query implies is what the bridge is trained for.
Or start from a worked example. The first five are held-out dataset frames with measured scores; the last three are stock aerial photos outside the training distribution.
Examples