Hugging Face has unveiled NeoMME, a multimodal encoder that processes text and images in a single bidirectional Transformer without separate vision towers or language models. The 260M and 800M parameter models represent a departure from industry norms, which typically adapt pretrained visual language models for retrieval tasks. NeoMME instead trains both modalities from scratch using a unified computational architecture, reducing parameter overhead while maintaining competitive performance on document retrieval benchmarks.
The model was trained using a masked discrete-diffusion objective on 524 billion tokens from multilingual text, code, mathematics, and images. Rather than preprocessing documents into separate visual and textual streams, NeoMME processes image patches and text tokens through the same Transformer encoder, enabling more efficient parallel processing and simplified deployment. The architecture uses dynamic image resolution to allocate computational resources based on content density, supports context lengths up to 16,384 tokens, and incorporates modern encoder techniques including grouped-query attention and alternating sliding-window and global-attention layers.
Benchmarked on visual document retrieval, the smaller 260M model achieves roughly 2x throughput compared to existing solutions, encoding approximately 51 pages per second on an NVIDIA L40S GPU. Quantization and hierarchical token pooling reduce storage requirements for embeddings by 255-fold—from 1.5 MB to 6 kB per page—while retaining over 95% of baseline performance. Hugging Face has released model checkpoints under an Apache 2.0 open-source license via Hugging Face Transformers.
Key Points
Unified architecture processes text and images without separate vision towers or pretrained components
Doubles throughput compared to competing models on document retrieval tasks
Reduces embedding storage by 255x through quantization while maintaining 95% baseline performance
Open-source under Apache 2.0 license and integrated into Hugging Face Transformers