Octen has introduced an AI-native web search API designed for agents, enhancing efficiency by breaking down questions into multiple parallel searches rather than processing them sequentially. This development addresses the limitations of existing web search systems, which are tailored for singular human queries, creating inefficiencies for AI agents that often conduct numerous simultaneous searches. The new API aims to improve performance metrics, as indicated by Octen’s published SealQA Hard benchmark results of 62ms for P50 and 68ms for P90, reflecting the infrastructure shift towards low-latency and high-concurrency performance in autonomous applications.

Octen: Octen is a search infrastructure company focused on building real-time, LLM-native web search APIs optimized for the generative AI era and agentic applications. Headquartered in San Francisco and Singapore, it develops distributed search engines designed to handle high-concurrency workloads for AI systems. The company recently launched an AI-native web search API that decomposes complex questions into multiple parallel searches to better support autonomous agent workflows.

Infrastructure Shift: New search layers are being engineered from the ground up to prioritize low-latency, high-concurrency performance for machine-to-machine interactions in autonomous applications.
AI Agent Optimization: Existing web search systems were primarily built for sequential human queries, creating a mismatch for AI agents that issue dozens of concurrent searches during research tasks.