Free Open Source Self Correcting-7B AI Model.
A Breakthrough in AI-Driven Deep Research
The new paper titled "PokeeResearch: Effective Deep Research via Reinforcement Learning from AI Feedback and Robust Reasoning Scaffold," introduces a groundbreaking 7B-parameter open-source AI agent designed to tackle complex research tasks with robustness and accuracy.
This work addresses critical limitations in current tool-augmented large language model, such as shallow retrieval, brittle tool-use, and weak alignment to factual correctness.
By leveraging reinforcement learning from AI feedback (RLAIF) and a sophisticated reasoning scaffold, PokeeResearch-7B sets a new standard for small-scale models in deep research, rivaling larger proprietary systems while remaining fully open-source.
The core of PokeeResearch-7B lies in its ability to decompose intricate queries, retrieve external evidence from tools like web searches, and synthesize grounded, verifiable responses.
Traditional AI agents often falter when tools fail or return noisy data, leading to hallucinations or incomplete answers. PokeeResearch overcomes these issues through an annotation-free RLAIF framework, where the model is trained using LLM-generated reward signals that evaluate factual accuracy, citation faithfulness, and adherence to user instructions. This self-improving loop allows the agent to optimize its policies without human annotations, making it scalable and efficient.
Complementing this is a chain-of-thought (CoT)-driven multi-call reasoning scaffold, which enables the agent to run multiple research threads in parallel, self-verify outputs for contradictions, and adaptively recover from errors. For instance, if a web tool returns irrelevant or erroneous information, the agent can pivot to alternative paths, ensuring resilient performance. The model's training emphasizes semantic correctness over superficial metrics like token overlap, allowing it to distinguish between plausible-sounding but incorrect responses and truly accurate ones.
Evaluated across 10 popular deep research benchmarks, PokeeResearch-7B demonstrates state-of-the-art results for models of its size. On challenging tasks like HLE (HotpotQA with Long Evidence), it achieves 17.6% accuracy; on GAIA (General AI Assistant benchmark), it scores 41.3%; and on BrowseComp (a web-browsing comprehension test), it reaches 8.4%. These figures surpass baselines like DeepResearcher by up to 17 points, highlighting the agent's superiority in handling real-world, multi-step research scenarios.
This not only advances the technical abilities of local AI but also democratizes powerful research tools, potentially accelerating progress toward more capable general AI systems.
I am running this model now.
The model is at
Paper:

24,05 tis.
38
Obsah na této stránce poskytují třetí strany. Není-li uvedeno jinak, společnost OKX není autorem těchto informací a nenárokuje si u těchto materiálů žádná autorská práva. Obsah je poskytován pouze pro informativní účely a nevyjadřuje názory společnosti OKX. Nejedná se o doporučení jakéhokoli druhu a nemělo by být považováno za investiční poradenství ani nabádání k nákupu nebo prodeji digitálních aktiv. Tam, kde se k poskytování souhrnů a dalších informací používá generativní AI, může být vygenerovaný obsah nepřesný nebo nekonzistentní. Další podrobnosti a informace naleznete v připojeném článku. Společnost OKX neodpovídá za obsah, jehož hostitelem jsou externí weby. Držená digitální aktiva, včetně stablecoinů a tokenů NFT, zahrnují vysokou míru rizika a mohou značně kolísat. Měli byste pečlivě zvážit, zde je pro vás obchodování s digitálními aktivy nebo jejich držení vhodné z hlediska vaší finanční situace.

