Awesome Multimodal Research

This repo is reorganized from Paul Liang’s repo: Reading List for Topics in Multimodal Machine Learning, feel free to raise pull requests!

News

[03/2023] OpenAI: ChatGPT plugins are tools designed specifically for language models with safety as a core principle, and help ChatGPT access up-to-date information, run computations, or use third-party services. https://openai.com/blog/chatgpt-plugins

“We’re also hosting two plugins ourselves, a web browser and code interpreter. We’ve also open-sourced the code for a knowledge base retrieval plugin, to be self-hosted by any developer with information with which they’d like to augment ChatGPT.”

[03/2023] Google Research: Bard is an early experiment that lets you collaborate with generative AI, powered by a research large language model (LLM), specifically a lightweight and optimized version of LaMDA. https://bard.google.com/

[03/2023] OpenAI: GPT-4 is a large multimodal model (accepting image and text inputs, emitting text outputs) that, while less capable than humans in many real-world scenarios, exhibits human-level performance on various professional and academic benchmarks. https://openai.com/research/gpt-4

[03/2023] Google Research: PaLM-E is a new generalist robotics model that overcomes these issues by transferring knowledge from varied visual and language domains to a robotics system. https://ai.googleblog.com/2023/03/palm-e-embodied-multimodal-language.html

[03/2023] OpenAI: ChatGPT and Whisper APIs, developers can now integrate ChatGPT and Whisper models into their apps and products through API. https://openai.com/blog/introducing-chatgpt-and-whisper-apis

[02/2023] MSR: Kosmos-1 is a multimodal large language model (MLLM) that is capable of perceiving multimodal input, following instructions, and performing in-context learning for not only language tasks but also multimodal tasks. https://github.com/microsoft/unilm#llm–mllm-multimodal-llm

[01/2023] Google Research: 2022 & beyond: Language, vision and generative models, a post of a series in which researchers across Google will highlight some exciting progress in 2022 and present the vision for 2023 and beyond. https://ai.googleblog.com/2023/01/google-research-2022-beyond-language.html

[11/2022] OpenAI: ChatGPT is a sibling model to InstructGPT, which is trained to follow an instruction in a prompt and provide a detailed response. https://openai.com/blog/chatgpt

[08/2022] MSR: Multimodal Pretraining: BEiT-3 is a general-purpose multimodal foundation model, which achieves state-of-the-art transfer performance on both vision and vision-language tasks. https://github.com/microsoft/unilm/tree/master/beit

[04/2022] OpenAI: DALL·E 2 is a new AI system that can create realistic images and art from a description in natural language. https://openai.com/dall-e-2/

[05/2021] Google: MuM, a new AI milestone for understanding information. https://blog.google/products/search/introducing-mum/

[03/2021] OpenAI: Multimodal Neurons in Artificial Neural Networks, which may explain CLIP’s accuracy in classifying surprising visual renditions of concepts, and is also an important step toward understanding the associations and biases that CLIP and similar models learn. https://openai.com/blog/multimodal-neurons/

[01/2021] OpenAI: CLIP maps images into categories described in text, and DALL-E creates new images from text. A step toward systems with deeper understanding of the world. https://openai.com/multimodal/