AllSpark Releases Iris-mini and Iris-pro: Open-Weight Search Agents Outperforming Proprietary Models on Key Benchmarks
AllSpark has announced the release of Iris-mini and Iris-pro, two new open-weight search-agent models. These models are designed to bridge the performance gap between proprietary and openly available search agents, offering advanced research capabilities without reliance on closed APIs. The release includes the models' weights on Hugging Face and their agent harness code on GitHub, making them accessible to researchers and developers.
Introducing Iris-mini and Iris-pro
The Iris series consists of two distinct models: Iris-mini, with 35 billion parameters, and Iris-pro, featuring 397 billion parameters. Both models are built upon the Qwen-series architecture and are engineered to operate with a substantial 256,000-token context window. This large context window is a critical component, as AllSpark's research suggests that effective runtime context management significantly influences benchmark performance, often more so than raw model quality.
Benchmark Performance and Training Methodology
AllSpark reports that Iris-mini and Iris-pro have achieved the strongest results among open-weight search agents within their respective size classes across four key benchmarks: BrowseComp, BrowseComp-ZH, DeepSearchQA, and Humanity's Last Exam. Specifically, Iris-mini scored 82.2, 84.8, 86.9, and 52.3 across these benchmarks, while Iris-pro achieved 88.6, 85.1, 92.9, and 56.4, respectively, with context management enabled.
The training pipeline for these models is innovative, involving the reverse-engineering of multi-step questions derived from web page link structures. These questions are then filtered and refined before undergoing reinforcement learning against live web search environments. This approach is designed to specialize the models for complex web search and long-horizon information-seeking tasks.
The Impact of Context Management
A notable finding from AllSpark's research is the significant role of runtime context management in enhancing model performance. The paper highlights that this aspect can often be the primary driver of performance differences observed in benchmarks. For instance, context management boosted Iris-mini's BrowseComp score by up to 21.2 points, demonstrating its critical impact on the model's ability to process and utilize information effectively during search tasks.
Availability and Future Implications
The weights for both Iris-mini and Iris-pro are publicly available on Hugging Face, allowing the broader AI community to access and experiment with these models. The accompanying agent harness code is also accessible on GitHub. This open-weight release is intended to foster further research and development in the field of conversational AI and autonomous agents, potentially accelerating advancements in AI news and applications.
The release of Iris-mini and Iris-pro represents a step towards democratizing access to advanced agentic search capabilities. By providing open-weight models that compete with or even surpass proprietary solutions in specific benchmarks, AllSpark aims to empower researchers and developers to build more sophisticated AI systems. The findings also underscore the importance of robust evaluation methods, as a case in the paper's appendix noted an agent's 'wrong' answer was actually correct against ground truth, suggesting a need for more nuanced benchmarks.
Sources
Recommended AI tools
ChatGPT
Conversational AI
AI research, productivity, and conversation—smarter thinking, deeper insights.
Perplexity
Search & Discovery
Clear answers from reliable sources, powered by AI.
Cursor
Code Assistance
Your coding agent for building ambitious software
DeepSeek
Conversational AI
Efficient open-weight AI models for advanced reasoning and research
Google Antigravity
Productivity & Collaboration
Build in the agent-first era
Grok
Conversational AI
Your cosmic AI guide for real-time discovery and creation
About the Author

Albert Schaper is a co-founder of Best-AI.org. He focuses on product strategy, AI adoption, practical tool selection, and educational content that helps users compare AI products with clearer context.
More from AlbertWas this article helpful?
Found outdated info or have suggestions? Send us a note.