任务域决策包
没有模型能可靠充当引用准确性的事实源
c
我需要综合 40 篇论文的研究发现
r
synthesize@document_analysis/literature_review
C'
跨论文论证映射用 Opus 4.8。批量单篇摘要用 Sonnet 4.6。所有引用都必须追溯到原文。
@document_analysis/literature_review
没有模型能可靠充当引用准确性的事实源
我需要综合 40 篇论文的研究发现
synthesize@document_analysis/literature_review
跨论文论证映射用 Opus 4.8。批量单篇摘要用 Sonnet 4.6。所有引用都必须追溯到原文。
cross-paper synthesis, identifying contradictions, thematic mapping
claude/opus-4.8batch summarization of individual papers, cost-sensitive
claude/sonnet-4.6any citation will appear in published work
human| ID | 关系 | 执行体 | 结果 | 输出 |
|---|---|---|---|---|
| CT-D5-001 | synthesize@cross_paper | claude/opus-4.8 | 已接受 | 最适合跨文档论证综合; maintains thread across sources |
| CT-D5-002 | synthesize@cross_paper | claude/sonnet-4.6 | 有边界 | strong on individual paper summaries; cross-paper coherence weaker |
| CT-D5-003 | validate@citation | openai/gpt-5.5 | 拒绝 | high hallucination risk on specific citations |
Opus 4.8 leads BrowseComp locating hard-to-find information online
Anthropic Opus 4.6 release / 4.8 confirmationAll current models require citation verification for literature review
根据相邻证据推断: hallucination data + general LLM behavior