MultimodalGoogle DeepMind Blog
Introducing Gemma 4 12B: a unified, encoder-free multimodal model
DeepMind has unveiled Gemma 4 12B, a multimodal model that combines text, images and other data types without separate encoders. With 12 billion parameters, the model aims to streamline architecture and boost efficiency on tasks that involve multiple modalities.
Summary written by Kernelia from the original article by Google DeepMind Blog. The story and its rights belong to its author.