Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
73 commits
Select commit Hold shift + click to select a range
d5e9e53
Add smart API asset policy for editable rebuilds
yiweiqin Jul 20, 2026
410acf6
Add global rebuild design planning
Sen-illion Jul 19, 2026
444fabc
Add raw OCR grouping for rebuild text
Sen-illion Jul 19, 2026
54e8da2
Add OCR text grouping
Sen-illion Jul 19, 2026
322b0b2
Add text layer ownership
Sen-illion Jul 19, 2026
7657a2c
Add deterministic rebuild visual quality report
Sen-illion Jul 19, 2026
a7bfac6
Add fallback rebuild preview renderer
Sen-illion Jul 19, 2026
9157b25
Tighten visual critic alignment grouping
Sen-illion Jul 19, 2026
e3bcbf9
Add paper-grounded editable figure workflow
yiweiqin Jul 20, 2026
91fd7ea
Restructure repository as a Codex plugin
yiweiqin Jul 20, 2026
879b9eb
Render editable card frames in rebuild outputs
yiweiqin Jul 20, 2026
0968cc6
Ignore local paper-to-editable smoke outputs
yiweiqin Jul 20, 2026
d09d12e
Introduce stable engine package boundaries
yiweiqin Jul 20, 2026
63ef1c4
Add dual scientific figure benchmark framework
yiweiqin Jul 20, 2026
09a933b
Add top-conference paper benchmark cases
yiweiqin Jul 21, 2026
2ef221d
Add structured PDF document extraction
yiweiqin Jul 21, 2026
18be9a8
Add PDF extraction quality gates and OCR fallback
yiweiqin Jul 21, 2026
497f2d7
Add fast framework prompt workflow
yiweiqin Jul 21, 2026
981c91f
Unify paper planning evidence contracts
yiweiqin Jul 21, 2026
9eb8838
Benchmark fast paper understanding on top-conference papers
yiweiqin Jul 22, 2026
f6892ab
Generalize and harden fast paper understanding
yiweiqin Jul 22, 2026
1309455
Document fast paper reliability workflow
yiweiqin Jul 22, 2026
e8c289c
Harden fast paper contracts across NLP architectures
yiweiqin Jul 22, 2026
a3b1a96
Improve OCR and long-paper evidence scheduling
yiweiqin Jul 22, 2026
8aac682
Handle rotated and multi-column paper layouts
yiweiqin Jul 22, 2026
0e1a81f
Speed up OCR and preserve sampled scan contracts
yiweiqin Jul 22, 2026
1d274a5
Recover two-column OCR layouts from title fragments
yiweiqin Jul 22, 2026
186f039
Support CJK papers without speculative relation repair
yiweiqin Jul 22, 2026
dce151c
Harden paper contracts against VLM schema drift
yiweiqin Jul 22, 2026
173a62e
Constrain paper contracts to evidence-connected graphs
yiweiqin Jul 23, 2026
bd12317
Recover unnumbered PDF headings from typography
yiweiqin Jul 23, 2026
66de30f
Add reproducible PDF extraction stress benchmarks
yiweiqin Jul 23, 2026
c01994c
Repair missing spaces in English OCR text
yiweiqin Jul 23, 2026
4725046
Profile RapidOCR extraction stages
yiweiqin Jul 23, 2026
fafc8a0
Avoid caching incomplete OCR runs
yiweiqin Jul 23, 2026
2065f56
Enforce OCR deadlines with isolated workers
yiweiqin Jul 23, 2026
cb95ffb
Bound single-page OCR by deadlines
yiweiqin Jul 23, 2026
d3ad3dc
Recover elided embedding lists from scanned papers
yiweiqin Jul 23, 2026
b4a00e9
Ground architecture context by local evidence
yiweiqin Jul 23, 2026
45ab79c
Recover high-value scan evidence within deadlines
yiweiqin Jul 23, 2026
90325b5
Add reproducible scanned paper benchmarks
yiweiqin Jul 23, 2026
e36b162
Bound fast VLM planning to the workflow deadline
yiweiqin Jul 23, 2026
88ccbc6
Filter repeated PDF margins and preserve rotated reading order
yiweiqin Jul 23, 2026
68538fc
Preserve multilingual labels and repair PDF hyphenation
yiweiqin Jul 23, 2026
bc7ea61
Recover numbered CJK sections from OCR
yiweiqin Jul 23, 2026
020414e
Enforce evidence for every visible contract entity
yiweiqin Jul 23, 2026
1add7f2
Handle rotated margins and repeated section titles
yiweiqin Jul 23, 2026
4dbe172
Fail cleanly on encrypted and damaged PDFs
yiweiqin Jul 23, 2026
be480ec
Benchmark real OCR on scanned and skewed columns
yiweiqin Jul 23, 2026
2e96d1b
Add OCR plausibility gates and structured table evidence
yiweiqin Jul 23, 2026
46bb280
Improve PDF extraction fidelity and contract grounding
yiweiqin Jul 23, 2026
29db4cd
Stabilize fast paper image generation
yiweiqin Jul 23, 2026
4c31b9d
Add evidence-aligned Image2 repair workflow
yiweiqin Jul 23, 2026
321acb1
Validate Image2 feedback-loop topology
yiweiqin Jul 23, 2026
745c040
Validate Image2 multi-head branch topology
yiweiqin Jul 23, 2026
0d7f628
Validate Image2 multimodal convergence topology
yiweiqin Jul 23, 2026
e501c1c
Validate Image2 dense multiframe topology
yiweiqin Jul 23, 2026
9c4b197
Add repeated Image2 stability audits
yiweiqin Jul 23, 2026
7521736
Stabilize feedback and dense Image2 topology
yiweiqin Jul 24, 2026
2f51ae6
Add generic paper semantic blueprint compiler
yiweiqin Jul 24, 2026
ac3e9a1
Compile paper contracts to editable PowerPoint
yiweiqin Jul 24, 2026
b614f68
Strengthen paper role and visual enrichment review
yiweiqin Jul 24, 2026
883506e
Normalize branch heads and feedback enrichment
yiweiqin Jul 24, 2026
c36be06
Stabilize dense overview normalization
yiweiqin Jul 24, 2026
0d9706a
Validate unseen paper contracts and PPT round trips
yiweiqin Jul 24, 2026
58cd7fc
Stabilize dense paper figures and editable resume
yiweiqin Jul 24, 2026
72d2d6f
Stabilize panel-local paper figure contracts
yiweiqin Jul 24, 2026
47bd591
Fix editable framework compatibility
yiweiqin Jul 24, 2026
622b6ea
Bridge fast paper contracts into editable frameworks
yiweiqin Jul 24, 2026
7fb9d23
Add evidence-grounded iterative image review
yiweiqin Jul 24, 2026
d3a84cb
Add evidence-aware benchmark re-auditing
yiweiqin Jul 25, 2026
be28bc3
Add deterministic labels and connector overlays
yiweiqin Jul 25, 2026
ed49f26
Normalize Image2 substrates for editable overlays
yiweiqin Jul 25, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
29 changes: 29 additions & 0 deletions .codex-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,29 @@
{
"name": "research-figure-studio",
"version": "0.2.0",
"description": "Generate paper-grounded scientific figures as editable PowerPoint compositions.",
"author": {
"name": "ResearchFigureStudio contributors",
"url": "https://github.com/yiweiqin/ResearchFigureStudio"
},
"homepage": "https://github.com/yiweiqin/ResearchFigureStudio",
"repository": "https://github.com/yiweiqin/ResearchFigureStudio",
"license": "MIT",
"keywords": ["research", "scientific-figures", "powerpoint", "pptx", "codex"],
"skills": "./skills/",
"interface": {
"displayName": "Research Figure Studio",
"shortDescription": "Paper to editable PowerPoint figures",
"longDescription": "Build scientifically grounded framework figures whose exact labels and relations remain editable in PowerPoint.",
"developerName": "ResearchFigureStudio contributors",
"category": "Productivity",
"capabilities": ["Read", "Write"],
"websiteURL": "https://github.com/yiweiqin/ResearchFigureStudio",
"defaultPrompt": [
"Turn this paper into an editable PPT framework figure.",
"Rebuild this scientific figure as editable PowerPoint.",
"Check whether this figure matches the paper semantics."
],
"brandColor": "#2F6F8F"
}
}
8 changes: 4 additions & 4 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -20,16 +20,16 @@ jobs:
- name: Compile package and bundled scripts
run: |
python -m compileall -q rfs
python -m py_compile codex-skills/research-figure-making/scripts/validate_framework_outputs.py
python -m py_compile codex-skills/research-figure-making/scripts/estimate_asset_fill.py
python -m py_compile skills/research-figure-studio/scripts/validate_framework_outputs.py
python -m py_compile skills/research-figure-studio/scripts/estimate_asset_fill.py
- name: Run unit tests
run: python -m unittest discover -s tests -q
- name: Check bundled skill metadata
shell: python
run: |
from pathlib import Path
skill = Path('codex-skills/research-figure-making/SKILL.md')
skill = Path('skills/research-figure-studio/SKILL.md')
text = skill.read_text(encoding='utf-8')
assert text.startswith('---'), 'SKILL.md must start with YAML frontmatter'
assert 'name: research-figure-making' in text
assert 'name: research-figure-studio' in text
assert 'description:' in text
5 changes: 5 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -20,6 +20,11 @@ research_figure_studio.egg-info/
venv/
env/
tmp/translation_env/
tmp/codex-presentations/
tmp/pdfs/*/
tmp/pdfs/
tmp/paper-to-editable-smoke*/
benchmarks/**/inputs/
*.key
*.pem

Expand Down
196 changes: 182 additions & 14 deletions README.md

Large diffs are not rendered by default.

70 changes: 70 additions & 0 deletions benchmarks/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,70 @@
# ResearchFigureStudio Benchmarks

ResearchFigureStudio uses two independent product benchmark suites plus a generated PDF extraction stress suite, so failures can be attributed to paper parsing, reference-image generation, or editable reconstruction.

## Suites

### `paper-to-image`

Measures scientific faithfulness, terminology, relation correctness, information coverage and density, clarity, aesthetics, reference compliance, hallucinations, and stability across repeated generations.

Scientific errors are hard failures and cannot be offset by aesthetics.

### `image-to-ppt`

Measures rendered visual fidelity, object and relation reconstruction, text alignment, editable PowerPoint structure, anti-cheating rules, visual blockers, and eventually edit-mutation behavior.

A full-slide copy of the reference image is a hard failure even if pixel similarity is perfect.

### Generated PDF extraction stress suite

`benchmark pdf-suite` creates deterministic native two-column, unnumbered-bold-section, rotated-page, and mixed native/scanned PDFs at runtime. It validates section boundaries, cross-column reading order, displayed coordinates, caption recovery, local OCR scheduling, English OCR spacing recovery, and elapsed time. `--ocr-engine auto` adds a real installed-OCR probe after the deterministic adapter-backed tier.

## Commands

```powershell
rfs benchmark list --root benchmarks --json
rfs benchmark validate --case benchmarks/paper-to-image/cases/001_linear_pipeline --json
rfs benchmark fetch --case benchmarks/paper-to-image/cases/101_vit_linear --json
rfs benchmark fast --case benchmarks/paper-to-image/cases/101_vit_linear --out output/benchmarks/vit_fast --planner-mode heuristic --json
rfs benchmark fast-suite --root benchmarks --out output/benchmarks/fast_suite --planner-mode heuristic --json
rfs benchmark pdf-suite --out output/benchmarks/pdf_extraction --ocr-engine auto --json
rfs benchmark run --case benchmarks/paper-to-image/cases/001_linear_pipeline --out output/benchmarks/p2i_001 --json
rfs benchmark score --case benchmarks/image-to-ppt/cases/001_three_stage_layout --run output/rebuild_case --json
```

## Case policy

Each case contains `case.json` plus suite-specific human-authored ground truth. Synthetic cases may be committed. Real papers and figures should only be committed when redistribution is permitted; otherwise use local paths or a private benchmark data repository.

Generated runs and reports belong under `output/benchmarks/` and are not source fixtures.

Real-paper cases commit `source.json` and human-authored Ground Truth, while `benchmark fetch` downloads the PDF into an ignored local `inputs/` directory. This keeps the public repository reproducible without redistributing publisher files.

The committed unseen-generalization set spans vision, multimodal learning, and NLP: ViT, Mask R-CNN, Self-Refine, ImageBind, SAM, DETR, CLIP, NeRF, Transformer, BERT, and retrieval-augmented generation. The NLP cases specifically exercise short overview captions, method-text recovery, embedding summation, pre-training/fine-tuning separation, and retrieval-conditioned generation. Local image-only scan regressions remain under ignored `tmp/pdfs/` and `output/pdf/`; do not commit publisher PDFs or generated OCR artifacts.

Any locally available PDF benchmark can also be tested as an image-only degradation without committing a publisher scan:

```powershell
rfs benchmark fast-suite `
--root benchmarks `
--out output/benchmarks/scanned `
--case-id 106_detr_set_prediction `
--case-id 110_bert_pretrain_finetune `
--rasterize-dpi 144 `
--ocr-engine rapidocr `
--planner-mode heuristic `
--deadline 180 `
--json
```

The generated image-only PDFs stay under the benchmark output directory. Score them with the same human-authored entity, relation, and forbidden-content contracts as the native PDFs.

`benchmark fast-suite` writes `fast_suite_report.json` with per-case results and aggregate entity/relation recall, forbidden content, cache hit rates, provider success/retry counts, failure categories, and stage timings.

## Evaluation tiers

- Offline contract tier: placeholder assets, deterministic validation, CI-safe.
- Fast planning tier: entity recall, relation recall, forbidden content, evidence grounding, deadline, and cache behavior without image generation.
- Production quality tier: real VLM/image models, frozen judge, repeated seeds, and human calibration.
- Human audit tier: blinded pairwise ratings for aesthetics, clarity, and information density.
33 changes: 33 additions & 0 deletions benchmarks/image-to-ppt/annotation_template.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,33 @@
{
"summary": "Independent human annotation template for image-to-PPT reconstruction.",
"case_id": "",
"canvas": {"width_px": 0, "height_px": 0},
"objects": [
{
"id": "object_01",
"type": "panel|card|asset|text|legend|formula",
"bbox_percent": {"x": 0.0, "y": 0.0, "w": 0.0, "h": 0.0},
"text": "",
"parent_id": null,
"z_index": 0,
"must_be_editable": true
}
],
"relations": [
{
"id": "relation_01",
"source": "object_01",
"target": "object_02",
"type": "data_flow",
"path_percent": [[0.0, 0.0], [0.0, 0.0]],
"must_be_editable": true
}
],
"groups": [],
"mutation_tests": [
{"operation": "replace_text", "target_id": "", "value": "Edited label"},
{"operation": "move_object", "target_id": "", "delta_percent": {"x": 0.05, "y": 0.0}},
{"operation": "change_fill", "target_id": "", "value": "#FFAA00"},
{"operation": "delete_relation", "target_id": ""}
]
}
26 changes: 26 additions & 0 deletions benchmarks/image-to-ppt/cases/001_three_stage_layout/case.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,26 @@
{
"summary": "Synthetic three-stage image-to-editable-PPT reconstruction case.",
"case_id": "001_three_stage_layout",
"suite": "image-to-ppt",
"tier": "offline-contract",
"reference_image": "reference.ppm",
"expected_objects": "expected_objects.json",
"preview": "rebuild_preview.png",
"pptx": "editable_composition.pptx",
"run_config": {
"asset_mode": "placeholder",
"asset_policy": "smart-api",
"text_mode": "off",
"layout_mode": "heuristic",
"control_mode": "heuristic",
"design_plan_mode": "heuristic",
"ocr_engine": "off"
},
"thresholds": {
"object_coverage": 0.8,
"relation_coverage": 1.0,
"editability_score": 0.66,
"full_slide_image_count": 0,
"blocking_visual_issue_count": 0
}
}
Original file line number Diff line number Diff line change
@@ -0,0 +1,14 @@
{
"summary": "Human-authored object and relation annotations for a three-stage layout.",
"objects": [
{"id": "input", "type": "card", "bbox_percent": {"x": 0.05, "y": 0.25, "w": 0.22, "h": 0.45}},
{"id": "method", "type": "card", "bbox_percent": {"x": 0.39, "y": 0.20, "w": 0.22, "h": 0.55}},
{"id": "output", "type": "card", "bbox_percent": {"x": 0.73, "y": 0.25, "w": 0.22, "h": 0.45}}
],
"relations": [
{"source": "input", "target": "method", "type": "data_flow"},
{"source": "method", "target": "output", "type": "data_flow"}
],
"required_editable_types": ["card", "text", "connector"],
"forbid_full_slide_reference_image": true
}
Original file line number Diff line number Diff line change
@@ -0,0 +1,6 @@
P3
6 3
255
250 250 250 250 250 250 250 250 250 250 250 250 250 250 250 250 250 250
80 150 210 80 150 210 250 250 250 70 180 160 70 180 160 220 150 80
80 150 210 80 150 210 250 250 250 70 180 160 70 180 160 220 150 80
26 changes: 26 additions & 0 deletions benchmarks/paper-to-image/cases/001_linear_pipeline/case.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,26 @@
{
"summary": "Synthetic three-stage pipeline for scientific and visual benchmark protocol validation.",
"case_id": "001_linear_pipeline",
"suite": "paper-to-image",
"tier": "offline-contract",
"paper": "paper.md",
"expected_semantics": "expected_semantics.json",
"preferences": "preferences.json",
"positive_references": [],
"negative_references": [],
"run_config": {
"planner_mode": "heuristic",
"image_asset_mode": "placeholder",
"image_candidates": 2,
"review_mode": "heuristic",
"aspect_ratio": "16:9",
"ocr_engine": "off"
},
"thresholds": {
"entity_recall": 1.0,
"relation_recall": 1.0,
"exact_label_rate": 1.0,
"hallucination_count": 0,
"forbidden_content_count": 0
}
}
Original file line number Diff line number Diff line change
@@ -0,0 +1,17 @@
{
"summary": "Human-authored semantic ground truth for the ExampleFlow figure.",
"entities": [
{"id": "raw_document", "label": "Raw Document", "role": "input", "required": true},
{"id": "document_encoding", "label": "Document Encoding", "role": "module", "required": true},
{"id": "evidence_aggregation", "label": "Evidence Aggregation", "role": "module", "required": true},
{"id": "final_prediction", "label": "Final Prediction", "role": "output", "required": true}
],
"relations": [
{"source": "raw_document", "target": "document_encoding", "type": "data_flow", "required": true},
{"source": "document_encoding", "target": "evidence_aggregation", "type": "data_flow", "required": true},
{"source": "evidence_aggregation", "target": "final_prediction", "type": "data_flow", "required": true}
],
"forbidden_labels": ["Agent", "Memory", "Retriever", "Knowledge Base"],
"required_innovation": "Evidence Aggregation",
"expected_reading_order": ["raw_document", "document_encoding", "evidence_aggregation", "final_prediction"]
}
11 changes: 11 additions & 0 deletions benchmarks/paper-to-image/cases/001_linear_pipeline/paper.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,11 @@
# ExampleFlow: A Three-Stage Evidence Processing Framework

## Abstract

ExampleFlow transforms a Raw Document into a Final Prediction through Document Encoding and Evidence Aggregation.

## Method

The system takes a Raw Document as input. Document Encoding converts the Raw Document into contextual representations. Evidence Aggregation consumes the contextual representations and produces a Final Prediction.

The processing order is Raw Document, Document Encoding, Evidence Aggregation, and Final Prediction.
Original file line number Diff line number Diff line change
@@ -0,0 +1,9 @@
{
"summary": "Visual preferences separated from scientific ground truth.",
"aspect_ratio": "16:9",
"language": "English",
"preferred_flow": "left_to_right",
"style_description": "clean academic framework figure",
"preferred_palette": ["blue", "teal", "warm orange accent"],
"avoid": ["commercial poster", "cyberpunk", "generic dashboard"]
}
15 changes: 15 additions & 0 deletions benchmarks/paper-to-image/cases/101_vit_linear/case.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,15 @@
{
"summary": "ICLR 2021 Vision Transformer case for a clean linear architecture figure.",
"case_id": "101_vit_linear",
"suite": "paper-to-image",
"tier": "real-paper-offline-contract",
"topology": "linear_pipeline",
"paper": "inputs/paper.pdf",
"source": "source.json",
"expected_semantics": "expected_semantics.json",
"preferences": "../../preferences/top_conference.json",
"positive_references": [],
"negative_references": [],
"run_config": {"planner_mode": "heuristic", "image_asset_mode": "placeholder", "image_candidates": 2, "review_mode": "heuristic", "aspect_ratio": "16:9", "ocr_engine": "off"},
"thresholds": {"entity_recall": 1.0, "relation_recall": 1.0, "exact_label_rate": 1.0, "plan_entity_recall": 0.8, "plan_relation_recall": 0.7, "hallucination_count": 0, "forbidden_content_count": 0, "plan_forbidden_content_count": 0}
}
Original file line number Diff line number Diff line change
@@ -0,0 +1,25 @@
{
"summary": "Human-authored semantic contract for the Vision Transformer architecture.",
"entities": [
{"id": "input_image", "label": "Input Image", "aliases": ["2D Image", "Image"], "role": "input", "required": true},
{"id": "image_patches", "label": "Image Patches", "aliases": ["Patches"], "role": "module", "required": true},
{"id": "linear_projection", "label": "Linear Projection of Flattened Patches", "aliases": ["Linear Projection"], "role": "module", "required": true},
{"id": "position_embedding", "label": "Position Embedding", "role": "module", "required": true},
{"id": "class_token", "label": "Class Token", "aliases": ["Class Embedding", "[class] Embedding"], "role": "module", "required": true},
{"id": "transformer_encoder", "label": "Transformer Encoder", "role": "module", "required": true},
{"id": "mlp_head", "label": "MLP Head", "role": "module", "required": true},
{"id": "class_prediction", "label": "Class Prediction", "aliases": ["Image Classification Predictions", "Classification Predictions", "Class Label"], "role": "output", "required": true}
],
"relations": [
{"source": "input_image", "target": "image_patches", "type": "data_flow", "required": true},
{"source": "image_patches", "target": "linear_projection", "type": "data_flow", "required": true},
{"source": "linear_projection", "target": "transformer_encoder", "type": "data_flow", "required": true},
{"source": "position_embedding", "target": "transformer_encoder", "type": "conditioning", "required": true},
{"source": "class_token", "target": "transformer_encoder", "type": "conditioning", "required": true},
{"source": "transformer_encoder", "target": "mlp_head", "type": "data_flow", "required": true},
{"source": "mlp_head", "target": "class_prediction", "type": "data_flow", "required": true}
],
"forbidden_labels": ["CNN Backbone", "Recurrent Network", "Decoder", "Object Detection Head"],
"expected_reading_order": ["input_image", "image_patches", "linear_projection", "transformer_encoder", "mlp_head", "class_prediction"],
"topology_requirements": {"primary_flow": "left_to_right", "branch_count": 0, "feedback_loop_count": 0, "dense_multi_panel": false}
}
12 changes: 12 additions & 0 deletions benchmarks/paper-to-image/cases/101_vit_linear/source.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,12 @@
{
"summary": "Publication metadata and local-fetch sources; the PDF is not redistributed by this repository.",
"paper_title": "An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale",
"venue": "ICLR",
"year": 2021,
"arxiv_id": "2010.11929",
"official_url": "https://openreview.net/forum?id=YicbFdNTTy",
"paper_urls": ["https://arxiv.org/pdf/2010.11929", "https://export.arxiv.org/pdf/2010.11929"],
"target_figure": "Figure 1",
"verified_caption": "Model overview. The image is split into fixed-size patches, linearly embedded, combined with position embeddings and processed by a Transformer encoder.",
"license_note": "Verify the source license before redistributing the PDF or figure image. Local benchmark fetching is for research use."
}
14 changes: 14 additions & 0 deletions benchmarks/paper-to-image/cases/102_mask_rcnn_branch/case.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,14 @@
{
"summary": "ICCV 2017 Mask R-CNN case for shared-backbone multi-head branching.",
"case_id": "102_mask_rcnn_branch",
"suite": "paper-to-image",
"tier": "real-paper-offline-contract",
"topology": "branching_architecture",
"paper": "inputs/paper.pdf",
"source": "source.json",
"expected_semantics": "expected_semantics.json",
"preferences": "../../preferences/top_conference.json",
"positive_references": [], "negative_references": [],
"run_config": {"planner_mode": "heuristic", "image_asset_mode": "placeholder", "image_candidates": 2, "review_mode": "heuristic", "aspect_ratio": "16:9", "ocr_engine": "off"},
"thresholds": {"entity_recall": 1.0, "relation_recall": 1.0, "exact_label_rate": 1.0, "plan_entity_recall": 0.8, "plan_relation_recall": 0.7, "hallucination_count": 0, "forbidden_content_count": 0, "plan_forbidden_content_count": 0}
}
Original file line number Diff line number Diff line change
@@ -0,0 +1,24 @@
{
"summary": "Human-authored semantic contract for Mask R-CNN.",
"entities": [
{"id": "input_image", "label": "Input Image", "aliases": ["Image"], "role": "input", "required": true},
{"id": "backbone", "label": "Backbone", "aliases": ["Backbone Architecture"], "role": "module", "required": true},
{"id": "region_proposals", "label": "Region Proposals", "aliases": ["Region Proposal Network (RPN)", "RPN", "Proposals"], "role": "module", "required": true},
{"id": "roi_align", "label": "RoIAlign", "role": "module", "required": true},
{"id": "classification_head", "label": "Classification", "aliases": ["Class Label", "Classification Head"], "role": "output", "required": true},
{"id": "box_head", "label": "Bounding-box Regression", "aliases": ["Bounding Box", "Box Branch", "Box Regression"], "role": "output", "required": true},
{"id": "mask_head", "label": "Mask Branch", "role": "output", "required": true}
],
"relations": [
{"source": "input_image", "target": "backbone", "type": "data_flow", "required": true},
{"source": "backbone", "target": "region_proposals", "type": "data_flow", "required": true},
{"source": "region_proposals", "target": "roi_align", "type": "data_flow", "required": true},
{"source": "backbone", "target": "roi_align", "type": "feature_flow", "required": true},
{"source": "roi_align", "target": "classification_head", "type": "branch", "required": true},
{"source": "roi_align", "target": "box_head", "type": "branch", "required": true},
{"source": "roi_align", "target": "mask_head", "type": "branch", "required": true}
],
"forbidden_labels": ["Semantic Segmentation Only", "Single Output Head", "Transformer Decoder"],
"expected_reading_order": ["input_image", "backbone", "region_proposals", "roi_align"],
"topology_requirements": {"primary_flow": "left_to_right", "branch_count": 3, "convergence_count": 1, "feedback_loop_count": 0}
}
Original file line number Diff line number Diff line change
@@ -0,0 +1,8 @@
{
"summary": "Publication metadata and local-fetch sources.",
"paper_title": "Mask R-CNN", "venue": "ICCV", "year": 2017, "arxiv_id": "1703.06870",
"official_url": "https://openaccess.thecvf.com/content_ICCV_2017/html/He_Mask_R-CNN_ICCV_2017_paper.html",
"paper_urls": ["https://arxiv.org/pdf/1703.06870", "https://export.arxiv.org/pdf/1703.06870"],
"target_figure": "Figure 1", "verified_caption": "The Mask R-CNN framework for instance segmentation.",
"license_note": "Verify redistribution rights; keep downloaded inputs local."
}
Loading
Loading