FoundationVision's Infinity Model Generates 1024px Images in 0.8 Seconds
FoundationVision has developed Infinity, a bitwise visual autoregressive text-to-image AI model capable of generating 1024×1024 photorealistic images from text prompts. The model uses a novel bitwise token prediction framework with an infinite-vocabulary tokenizer, allowing it to scale beyond the limitations of traditional autoregressive approaches. Infinity produces high-resolution images in 0.8 seconds without additional optimization, making it 2.6 times faster than SD3-Medium while outperforming it on key benchmarks including a GenEval score of 0.73 and an ImageReward score of 0.96. The model also achieved a 66% human preference win rate over competing diffusion models such as SD3-Medium and SDXL. Built on PyTorch, Infinity's weights are publicly available on Hugging Face, making it accessible for use cases ranging from e-commerce imagery to real-time design tools.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in