I keep hearing that future models will be inherently multimodal. Does this mean they will handle audio, video, and text in one single pass? If I am building a tool right now, should I be focusing on models that do this, or are specialized models for each modality still going to perform better? I am trying to future-proof my architecture, but it feels like the goalposts are shifting every week.
The question