AI
The module that sits between the camera and the language model
A newly published application describes a video-based multimodal system with a third component wedged between the image encoder and the language model. The placement is the whole invention.