Modality bridge
Make language, vision, video, structure, and action mutually legible.
A language-first video learner that converts frames into structured image-and-language descriptors for few-shot video tasks.
Diagnosed the action-knowledge gap in video-language models and patched temporal action dynamics into frozen foundation models.
Introduced Primal Visual Description, a symbolic bridge from precise vector perception to language reasoning.
- DyMU ↗NeurIPS 2025
Dynamically merges—and virtually restores—visual tokens to accelerate VLM inference without permanently discarding information.


