Multimodal AI

Systems that connect visual, textual, and other modalities, with a focus on multimodal understanding and reasoning.

Overview

A detailed overview is being prepared.

Open questions

Open questions for this area are being developed.

Current projects

Our first projects in this area will appear here.

Publications

Publications in this area are in progress.

Researchers

Researcher profiles for this area are being prepared.

Resources

Datasets and code for this area will appear here.

Other areas in Junction

  1. Vision-Language Models