Hacker News
- Rien ce jour-là.
Over the past few years, multimodal large language models have become increasingly capable of understanding images, videos, and real-world scenes. They can recognize objects, reason about spatial relationships, answer visual questions, and solve complex multimodal reasoning tasks. But for embodied intelligence, understanding the world is only the first step. A truly embodied agent also needs to un