Y. Wang
University of Illinois Urbana-Champaign, Illinois, United States
Keywords: Long-video understanding, Self-improving embodied agents, Autonomous systems, Streaming perception, Humanoid loco-manipulation
Long-video understanding is emerging as a critical capability for autonomous agents operating in complex real-world environments, where intelligence depends on extracting actionable information from massive streams of visual experience. However, scaling to hours or days of video requires systems that can efficiently preserve mission-relevant information while discarding redundancy. This work presents a unified framework for dynamic and learnable context compression for long-horizon video understanding. Our approach combines scalable map-reduce architectures for agentic video reasoning, supervised fine-tuning for extreme token compression, and reinforcement learning for precise localization of key events and decision-critical moments. We further extend these capabilities to streaming and interactive environments, where adaptive memory management enables persistent perception and long-term situational awareness under compute and bandwidth constraints. Beyond passive understanding, we demonstrate how agents can leverage extended visual experience to autonomously acquire new behaviors and skills, enabling self-improving embodied systems for real-world operation. These capabilities support dual-use civilian and defense applications including intelligent AI assistants, autonomous surveillance, human-robot teaming, logistics automation, and humanoid loco-manipulation in dynamic environments. This work advances scalable visual intelligence for persistent autonomous systems operating in long-duration real-world missions.