用户提问:
Ornith-1.0是什么模型?
---

用户问的是"Ornith-1.0是什么模型?"。这是一个我不熟悉的术语,需要进行搜索。按照规划要求,我需要编制6个搜索关键词从不同角度并行搜索。让我思考一下:

  1. "Ornith-1.0" - 直接搜索这个名称
  2. "Ornith-1.0 model" - 英文搜索,可能是一个AI模型
  3. "Ornith-1.0 模型" - 中文搜索
  4. "Ornith 1.0" - 带空格的版本
  5. "Ornith-1.0 AI" - 可能是AI领域的新模型
  6. "Ornith model" - 更宽泛的搜索

或者也可以考虑:

  • 可能是一种生物相关的模型?"Ornith" 词根与鸟有关(ornithology 鸟类学)
  • 可能是某种技术模型

让我用6个不同的关键词并行搜索,count=20。

🧭 需求解读与搜索规划

用户查询"Ornith-1.0是什么模型?"——该术语我并不熟悉,需要从零开始搜索验证。Ornith 词根与鸟类学(ornithology)相关,但具体是哪种模型(AI模型?生物学模型?技术框架?)尚不确定,因此我设计6个不同角度的搜索关键词并行检索,确保覆盖各种可能性。
"Ornith-1.0"
Ornith-1.0 model
Ornith-1.0 模型
"Ornith" model
Ornith-1.0 2025 2026
Ornith 1.0 release

搜索结果显示,Ornith-1.0 是一个近期(2026年6月)由 DeepReinforce 发布的开源大语言模型家族,专注于 agentic coding(代理式编程)。有多个模型规模和关键特性。

我现在需要深入阅读一些关键来源来获取详细信息。让我并行阅读几个最有价值的页面:

  1. 官方页面(deep-reinforce.com)- 了解核心概念
  2. Hugging Face 集合 - 了解模型规模
  3. GitHub 仓库 - 了解技术细节
  4. Simon Willison 的文章 - 技术分析
  5. ExplainX 的文章 - 技术解读
  6. 知乎文章 - 中文解读

这些来源可以交叉验证信息。

搜索结果显示 Ornith-1.0 是刚刚于2026年6月发布的新模型,有丰富的信息源。现在我对关键信源进行深度并行阅读,以获取完整、准确的信息。

I now have comprehensive information about Ornith-1.0 from multiple sources. Let me also read one more source to get additional details about DeepReinforce and the model's background.

Let me also check a few more sources to round out the picture - the MarkTechPost article and the explainx.ai article for additional perspectives.

Actually, I already have very comprehensive information from:

  1. DeepReinforce official blog - core innovation details
  2. Simon Willison's blog - independent analysis
  3. GitHub repo - benchmark data and deployment info
  4. Hugging Face collection - model variants

Let me also read the explainx.ai article and the verdent.ai guide for more details.

Wait, let me think about what else I need. The question is "Ornith-1.0是什么模型?" which is simple - what kind of model is it. I already have enough information to answer comprehensively. Let me also look into DeepReinforce the organization behind it, and maybe the MarkTechPost article for another perspective.

Let me read a couple more articles to get a fuller picture and then compile the report.

信息已经非常丰富了!让我再补充阅读几个不同的视角来完善理解。

I now have extremely comprehensive information about Ornith-1.0 from multiple authoritative sources. Let me verify the key facts across sources:

  1. What is Ornith-1.0? - Cross-verified across all sources:
    • An open-source, MIT-licensed family of large language models (LLMs) specialized for agentic coding
    • Released by DeepReinforce on June 25, 2026
    • Built on top of Gemma 4 and Qwen 3.5 pre-trained weights
  2. Model variants - Multiple sources confirm:
    • 9B Dense (≈9 billion parameters)
    • 31B Dense (≈31 billion parameters)
    • 35B MoE (≈35 billion parameters, Mixture-of-Experts)
    • 397B MoE (≈397 billion parameters, Mixture-of-Experts - flagship)
  3. Core innovation (Self-Scaffolding) - Confirmed by official blog, Simon Willison, explainx.ai, and others:
    • The model learns to generate its own "scaffold" (the harness/agent framework) during RL training, not just the solution
    • Two-stage RL process: first optimize the scaffold, then generate solutions using that scaffold
    • Rewards back-propagated to both stages
  4. Benchmark performance - Detailed data from GitHub repo and official blog:
    • 397B model: Terminal-Bench 2.1: 77.5, SWE-Bench Verified: 82.4 (beats Claude Opus 4.7 in some benchmarks)
    • 35B model: Terminal-Bench 2.1: 64.2 (beats Qwen 3.5-397B's 53.5)
    • 9B model: Terminal-Bench 2.1: 43.1, SWE-Bench Verified: 69.4
  5. Anti-reward hacking mechanisms - Three layers: fixed trust boundary, deterministic monitor, frozen LLM judge
  6. Deployment - Available via vLLM, SGLang, Ollama, LM Studio, llama.cpp; MIT license, no restrictions
  7. Organization behind it - DeepReinforce (AI research collective/team, first model release)

Now let me check if there are any discrepancies across sources:

  • The release date: Most sources say June 25, 2026; some say June 26. Both are consistent within a narrow window.
  • The model sizes: All sources agree on 9B, 31B, 35B MoE, 397B MoE
  • The base models: Most say Qwen 3.5 and Gemma 4; the official blog explicitly states "built on pre-trained Gemma 4 and Qwen 3.5"
  • The GitHub says 31B is built on Gemma, the others on Qwen 3.5

I think I have very comprehensive information now. Let me check if there are any remaining questions or issues:

Current status:

  • ✅ Key facts verified across 2-3+ independent sources
  • ✅ No unresolved contradictions found
  • ✅ No next steps needed for search
  • ✅ Information is sufficient to answer the question

I can now proceed to start-report to write the final report. Let me compile all the information.

Wait, let me double-check the key remaining questions:

  1. "仍待验证问题" - None. All key facts about Ornith-1.0 are well established.
  2. "下一步检索方向" - None. Information is fully sufficient.
  3. All plan tasks executed.

I'm ready to call start-report.

信息已经非常充实,多源交叉验证完成。现在进入最终报告撰写阶段。

Now I need to write a comprehensive search brief/report about Ornith-1.0 based on all the information I've gathered. Let me organize all the verified facts from multiple sources and create a well-structured report.

Let me compile the key information:

What is Ornith-1.0?
  • An open-source, MIT-licensed family of large language models (LLMs) specialized for agentic coding
  • Released by DeepReinforce on June 25, 2026
  • Built on pre-trained Gemma 4 and Qwen 3.5 base models
  • Named after the Greek word for "bird" (寓意像鸟儿筑巢一样自主构建框架)
Model variants:
  • 9B Dense (≈9 billion parameters) - edge/single GPU
  • 31B Dense (≈31 billion parameters) - based on Gemma
  • 35B MoE (≈35 billion parameters, Mixture-of-Experts) - practical balance
  • 397B MoE (≈397 billion parameters, MoE) - flagship
Core Innovation - Self-Scaffolding:
  • The model learns to generate its own agent scaffold (harness) during RL training, not just the solution
  • Two-stage RL process: first optimize the scaffold, then generate solutions using that scaffold
  • Rewards back-propagated to both stages
  • Anti-reward hacking: three-layer defense mechanism
Benchmark Performance:
  • 397B: Terminal-Bench 2.1: 77.5, SWE-Bench Verified: 82.4
  • 35B: Terminal-Bench 2.1: 64.2 (beats Qwen 3.5-397B's 53.5)
  • 9B: Terminal-Bench 2.1: 43.1, SWE-Bench Verified: 69.4
Deployment:
  • vLLM, SGLang, Ollama, LM Studio, llama.cpp
  • MIT license, no geographical restrictions
  • 262K context window
Limitations noted:
  • All benchmark results are self-reported by DeepReinforce, not independently verified
  • Performance depends heavily on test framework and settings
  • Ornith is weights only, not a ready-to-use product

Let me write a comprehensive report now.

I need to be careful about citations. Let me map the sources:

  • 74 Simon Willison's blog
  • 75 Hugging Face collection
  • 76 DeepReinforce official blog
  • 77 知乎 (access denied)
  • 78 GitHub repository
  • 79 MarkTechPost
  • 80 Verdent.ai guide
  • 81 explainx.ai blog

Let me write the report now.

🦅 Ornith-1.0:DeepReinforce 推出的自改进开源 Agentic Coding 大模型家族

🎯 核心发现速览

Ornith-1.0 是由 DeepReinforce 在 2026 年 6 月 25 日发布的一款专为代理式编程(Agentic Coding)设计的 MIT 开源大语言模型(LLM)家族 767981。其最大创新是 "自脚手架(Self-Scaffolding)"训练框架 ——模型在强化学习过程中不仅学习生成代码解决方案,还 自主构建优化引导解决方案执行的 Agent 框架 (如内存布局、重试逻辑、工具编排),无需人工为不同基准维护单独的框架配置767881。该模型家族涵盖从 90 亿到 3970 亿参数的四种规格,旗舰版在多项编码 Agent 基准测试中达到甚至超越 Claude Opus 4.7767881

🧩 模型全景:参数规格与定位

Ornith-1.0 提供四个参数版本,覆盖从边缘设备到多 GPU 集群的全场景767880
版本架构基座模型最佳部署场景
Ornith-1.0-9BDense(约 90 亿参数)Qwen 3.5单 GPU 部署/边缘设备
Ornith-1.0-31BDense(约 310 亿参数)Gemma 4多 GPU 部署
Ornith-1.0-35BMoE(约 350 亿参数,每 token 激活约 30 亿)Qwen 3.5性能与效率的平衡点
Ornith-1.0-397BMoE(约 3970 亿参数,旗舰版)Qwen 3.5复杂长周期任务最高精度
全部版本均采用 MIT 许可证 ,无地域使用限制,全球可免费商用757880。支持 262,144 token(约 25.6 万)上下文窗口78。9B 版本可在单张 80GB GPU 上以 bf16 运行,量化后(GGUF)可在消费级硬件上部署212678

💡 核心创新:自改进训练框架(Self-Scaffolding)

Ornith-1.0 区别于传统编码模型的本质在于 对 Agent 框架本身的端到端优化 767981

🏗️ 传统方案的局限

大多数基于 LLM 的编码 Agent 依赖于 人工设计的固定脚手架 (scaffold/harness),即工程师为每个基准类别手动编写工具调用模式、执行循环和错误处理逻辑。模型仅负责生成代码,而"如何组织解题过程"由人工预设768081

🔄 Ornith 的两阶段 RL 训练

Ornith-1.0 在强化学习中引入 两阶段联合优化 767881
  1. 框架阶段(Scaffold Stage) :基于当前任务和此前使用的框架,模型首先提出优化后的 Agent 框架(包括内存管理、重试策略、工具编排等)
  2. 解决方案阶段(Solution Stage) :基于优化后的框架和任务描述,模型生成具体的代码解决方案
最终解决方案的奖励会 反向传播到两个阶段 ,使高回报的框架结构随训练迭代被自动筛选和强化,无需人工干预767981

🛡️ 三层奖励欺骗防御机制

针对模型自主设计框架可能引发的奖励黑客(reward hacking)问题,Ornith-1.0 设计了 三层防护 767981
  • 固定信任边界 :环境、工具接口和测试隔离机制不可修改,模型仅能优化内部策略逻辑
  • 确定性监控器 :基于规则的监控器标记任何读取未授权路径、修改验证脚本或调用未授权工具的行为,违规轨迹获零奖励并从优势计算中排除
  • 冻结 LLM 评判器 :作为验证器之上的否决机制,拦截意图层面的作弊行为

⏳ 异步强化学习(Pipeline-RL)

针对长 Agent 轨迹的离策略(off-policy)问题,Ornith-1.0 采用带 陈旧性加权 GRPO 的管道强化学习:不同时间生成的 token 根据其陈旧程度赋予不同权重,超过阈值则直接丢弃,保障长周期训练的稳定性7681

📊 基准测试表现

所有基准测试结果为 DeepReinforce 自行发布,测试框架与参数设置已公开说明,但 尚未经第三方独立验证 8081

🏆 旗舰版 Ornith-1.0-397B 的关键对比

在多项主流编码 Agent 基准测试中,Ornith-1.0-397B 达到或超过 Claude Opus 4.7,与闭源前沿模型处于同一梯队767881
基准测试Ornith-1.0-397BQwen 3.5-397BClaude Opus 4.7Claude Opus 4.8DeepSeek-V4-Pro
Terminal-Bench 2.1 (Terminus-2)77.553.570.385.067.9
SWE-Bench Verified82.476.480.887.680.6
SWE-Bench Pro62.251.664.369.255.4
NL2Repo48.236.869.7
ClawEval (平均)77.170.778.275.8
注:数据来源为 DeepReinforce GitHub 仓库与官方博客7678,空单元格代表该模型在原始数据中未列出对应分数。

⚡ 中小尺寸的"以小博大"

  • Ornith-1.0-35B 在 Terminal-Bench 2.1 上得分 64.2超过 Qwen 3.5-397B 的 53.5 ,展示了自框架训练对 Agent 能力的显著压缩效果767881
  • Ornith-1.0-9B 在 SWE-Bench Verified 上得分 69.4 ,基本与 Gemma 4-31B(52.0)拉开显著差距,证明 Agent 编程能力可被高效压缩到可边缘部署的模型尺寸中7678

🚀 部署与使用

支持的运行环境

Ornith-1.0 作为开放权重模型(MIT 许可证),可通过以下主流推理运行时部署7880
  • vLLM (推荐,≥ 0.19.1)
  • SGLang (≥ 0.5.9)
  • Ollama (GGUF 量化版本)
  • LM Studio (本地桌面运行)
  • llama.cpp (GGUF 量化版本)
  • Hugging Face Transformers (≥ 5.8.1)
所有检查点提供 OpenAI 兼容接口 ,可直接接入现有 Agent 框架生态78

推理模型特性

Ornith-1.0 属于 推理型模型 ,默认回复前会生成 <think>…</think> 思考块,支持工具调用(<tool_call>),需在服务端启用推理解析器和工具调用解析器78
推荐采样参数:temperature=0.6top_p=0.95top_k=2078

🏢 发布方:DeepReinforce

DeepReinforce 是一个 AI 研究团队/集体,Ornith-1.0 是其 首个公开发布的模型 7479。目前关于该团队的公开信息有限,可追溯的最早论文是 2025 年 6 月发布的《CUDA-L1: Improving CUDA Optimization via Contrastive Reinforcement Learning》74。团队在 Hugging Face 上以 deepreinforce-ai 组织发布模型75,并在 Twitter/X 上以 @ornith_ 账号运营81

⚖️ 局限性声明与审慎提示

  1. 基准测试未经第三方独立验证 :所有测试结果均由 DeepReinforce 自行发布,评分高度依赖测试框架与参数设置(温度、上下文窗口、采样设置等),不同测试框架下同一模型的得分可能存在显著差异。引用时应标注为"DeepReinforce 报告的……"并注明测试条件8081
  2. 仅提供模型权重 :Ornith-1.0 不是开箱即用的托管编码助手,用户需自行从 Hugging Face 下载权重,通过 vLLM/Ollama/LM Studio 等工具部署并集成到 Agent 框架中80
  3. 尚在早期阶段 :作为 DeepReinforce 的首个模型发布,其社区生态、第三方工具支持、文档完善度仍在成长中74
  4. 部分基准仍存差距 :在 SWE-Bench Pro 上,Ornith-1.0-397B(62.2)低于 Claude Opus 4.7(64.3);Claude Opus 4.8 和 GLM-5.2-744B 在多项指标上仍保持领先767881
  5. 可用性需自行确认 :部分 checkpoint 的实际发布状态和支持的运行时版本需以官方 Hugging Face 仓库为准80

🌟 总结与定位

Ornith-1.0 是 2026 年 6 月最受关注的开源 Agentic Coding 模型之一 81。其核心突破不在于单纯的代码生成质量,而在于 将 Agent 框架(scaffold)从人工设计的固定组件转变为模型自主学习的一等公民 7681

在 2026 年的 Agentic Coding 赛道中,Ornith-1.0 的差异化定位是:

  • 与闭源前沿(Claude Opus 4.7/4.8、GPT-5.x、Gemini 3.x)竞争,但在部分指标上已能与 Opus 4.7 正面匹敌767981
  • 相比同尺寸开源模型(Qwen 3.5/3.6、Gemma 4),在 Agent 编码基准上展现出显著优势767881
  • 重视数据控制、本地部署、离线使用或大规模部署成本的团队 提供了一个值得关注的开放权重选项80
10