Voxae

Language-grounded, metric scene understanding for physical AI.

Ask an aerial scene a question in plain language and get the region that answers it. Two systems run on the same query:

  • Trained <SEG> bridge: a vision-language model whose hidden state at a <SEG> token is projected directly into SAM 2.1's prompt space and trained end to end on Voxae-Reason (referring + affordance queries over drone imagery).
  • Zero-shot baseline: a hosted VLM emits a box and points as text, SAM 2.1 turns them into a mask. No training.

The gap is widest on implicit queries: naming an object is easy, reasoning about what a query implies is what the bridge is trained for.

Or start from a worked example. The first five are held-out dataset frames with measured scores; the last three are stock aerial photos outside the training distribution.

Examples