OCG LabTask-Domain Decision System
Search evidence
Live
中文
D1bounded

Code Engineering / Single-File Debug

@code_engineering/single_file_debug

5 transitions4 decisions3 evidence
cost-benefit threshold at ~300 lines / single-call scope
D2rejected

Code Engineering / Large Repo Migration

@code_engineering/large_repo_migration

4 transitions4 decisions3 evidence
file-level decomposition is the only stable path for cross-file state tasks
D3rejected

Code Engineering / Code Review

@code_engineering/code_review

4 transitions4 decisions2 evidence
SecPass 19% across all models; security review always needs human gate
D4rejected

Document Analysis / Long PDF Extraction

@document_analysis/long_pdf_extraction

3 transitions4 decisions3 evidence
multi-needle retrieval is the capability cliff for Sonnet 4.6
D5rejected

Document Analysis / Literature Review

@document_analysis/literature_review

3 transitions3 decisions2 evidence
no model is a reliable source of truth for citation accuracy
D6rejected

Structured Data / JSON Extraction

@structured_data/json_extraction

3 transitions4 decisions2 evidence
schema validation layer is non-negotiable regardless of model
D7bounded

Structured Data / SQL Repair

@structured_data/sql_repair

3 transitions3 decisions2 evidence
engine-specific optimization is the capability threshold between Sonnet and Opus
D8bounded

Language Reasoning / Legal Clause Comparison

@language_reasoning/legal_clause_comparison

4 transitions3 decisions4 evidence
no model is authoritative on binding legal decisions; always human gate
D9bounded

Business Decision / Customer Support Triage

@business_decision/customer_support_triage

3 transitions4 decisions3 evidence
Opus 4.8 cost unjustified for routine triage; Sonnet 4.6 is the correct default
D10rejected

Structured Data / Financial Table Extraction

@structured_data/financial_table_extraction

3 transitions4 decisions4 evidence
numerical verification against source is mandatory regardless of model