ResearchFigureStudio is a PPTX-first research figure generation pipeline for paper-grounded, reference-guided scientific framework figures. It turns a paper plus a user-provided visual reference image into an editable PowerPoint composition assembled from many slot-level image assets, not one flattened full-diagram bitmap.
The current workflow is optimized for AI/ML/NLP system figures:
- paper-grounded concept extraction
- reference-primary geometry, style, color, and flow alignment
- 25-50 non-arrow image slots
slot_visual_spec.jsonfor dense mini-scene/image-block planning- AutoFigure-inspired control candidates and overlays for arrow/source-target binding
- reference-preserving arrow styling/routing reports for softer editable PPT connectors
- reference-constrained orthogonal fallback routing for missing or explicitly fallback-allowed connectors
- optional
--arrow-style-mode aestheticfor reference-tunnel arrow beautification with curve connectors, halo underlays, and explicit opt-in bundle lane offsets - multi-candidate image generation through placeholder, Gemini, or Yunwu image2-compatible APIs
- deterministic PPTX composition with editable labels, panels, arrows, connectors, and formulas
- optional Presentations-plugin QA for importing/rendering/inspecting the PPTX without mutating it
- strict validation for no single full diagram, no semantic crop, no vector-only fallback, low blank space, and non-trivial image-block complexity
This repository does not include API keys, papers, reference images, generated outputs, or local run artifacts.
ResearchFigureStudio is an early engineering prototype. Its current practical ability is limited to placing many small generated image blocks into precise PowerPoint positions, then keeping surrounding labels, arrows, panels, formulas, and grouping elements editable in PPTX.
It does not yet solve the harder goal of generating truly publication-grade, fully editable scientific figures end to end. The image blocks themselves are still raster assets, not editable scientific vector objects. The system can produce a PPTX composition that is easier to manually revise, but it should not be treated as a finished top-tier-paper figure generator.
The most important open problem is reliable arrow and connector localization.
The current implementation now has an initial AutoFigure-inspired
reference_control_candidates.json plus slot_overlay.png /
reference_control_overlay.png workflow, plus reference-preserving
arrow_style_profile.json, selected_arrow_routes.json, and
arrow_quality_report.json. It now includes a conservative orthogonal fallback
router for missing or explicitly fallback-allowed connectors. It is still
fragile for complex scientific diagrams: source-target binding, truly curved
routes, dense bundle routing, dashed loops, and preserving reference-image logic
need stronger methods.
If you have experience with vision-language layout parsing, diagram structure reconstruction, PowerPoint object routing, graph drawing, or editable scientific figure generation, guidance and contributions are very welcome.
See CONTRIBUTING.md and docs/help-wanted.md for concrete contribution areas.
For GPT/Codex agents reproducing or continuing this workflow, start with docs/gpt-reproduction-workflow.md.
git clone https://github.com/yiweiqin/ResearchFigureStudio.git
cd ResearchFigureStudio
python -m pip install --upgrade pip
python -m pip install -e .
rfs doctor --jsonOptional local OCR support for reference-derived editable text:
python -m pip install -e ".[ocr]"Windows with PowerPoint installed gives the best PPTX/PDF/PNG export path. The Python package itself can still generate and validate most intermediate artifacts without external image APIs when using --asset-mode placeholder.
Use paper-to-image when the required endpoint is a generated raster framework
figure rather than an editable PowerPoint file. The production route performs a
universal evidence-grounded paper review, loads a domain extension, converts
positive references into content-free architecture templates, selects a template,
renders layout_blueprint.png, and uses Image2 edit to create and review three
candidates. It never invokes the PPTX compiler.
Offline engineering validation:
rfs paper-to-image `
--paper "C:\path\paper.pdf" `
--out "output\paper_to_image_placeholder" `
--planner-mode heuristic `
--asset-mode placeholder `
--candidates 2 `
--review-mode heuristic `
--ocr-engine off `
--jsonPlaceholder output is written as engineering_preview.png. It is explicitly
ineligible for production delivery and never becomes selected_image.png.
Real VLM planning and Image2 generation:
rfs paper-to-image `
--paper "C:\path\paper.pdf" `
--out "output\paper_to_image" `
--planner-mode vlm `
--domain-profile auto `
--positive-reference "C:\path\reference1.png" `
--positive-reference "C:\path\reference2.png" `
--template auto `
--asset-mode image2 `
--candidates 3 `
--aspect-ratio auto `
--review-mode vlm `
--repair-rounds 1 `
--ocr-engine auto `
--jsonThe main outputs are paper_review.json, review_coverage_report.json,
domain_profile.json, template_profiles/, selected_template.json,
layout_blueprint.png, figure_specification.json, image_prompt.txt,
image2_request_manifest.json, four critic reports, the candidates/
directory, and production-only selected_image.png.
Reference-conditioned production generation requires an Image2 edit endpoint.
It defaults to <API_BASE>/images/edits and may be overridden with
RFS_IMAGE_EDIT_URL. No API key value is written to output artifacts or logs.
If a key was pasted into a chat, issue, or terminal transcript, revoke it and use
a newly rotated key through environment variables before running production.
See docs/paper-to-image.md for the review schema,
template contract, production gates, and failure behavior.
Use placeholder assets to validate the local pipeline without calling any API:
rfs make-framework `
--paper "C:\path\paper.pdf" `
--reference "C:\path\reference.png" `
--out "output\demo_placeholder" `
--slot-count 25 `
--slot-source reference-primary `
--complexity-profile reference-dense `
--candidates-per-slot 2 `
--locator-mode heuristic `
--control-localizer-mode heuristic `
--arrow-style-mode reference `
--prompt-plan-mode heuristic `
--asset-mode placeholder `
--asset-workers 4 `
--asset-review-mode heuristic `
--critic-mode heuristic `
--text-extractor-mode ocr `
--ocr-engine paddle `
--ocr-lang en_ch `
--json
rfs validate --out "output\demo_placeholder" --jsonoutput/ is intentionally ignored by Git.
Set API credentials only through environment variables. Do not write keys into source files.
$env:API_BASE='https://yunwu.ai/v1'
$env:API_KEY='<your key>'
$env:GEMINI_API_KEY=$env:API_KEY
$env:GEMINI_GEN_IMG_URL='https://yunwu.ai/v1beta/models/gemini-2.5-flash-image:generateContent'
$env:MODEL_VLM='gemini-3-pro-preview-thinking'
$env:RFS_PROMPT_PLANNER_MODEL=$env:MODEL_VLM
$env:RFS_CONTROL_LOCALIZER_MODEL=$env:MODEL_VLM
$env:RFS_IMAGE_MODEL='image-2'For the reference-only image-to-editable-PPT workflow, configure the same VLM credentials plus the image-generation endpoint:
$env:API_BASE='https://your-openai-compatible-provider/v1'
$env:API_KEY='<your key>'
$env:MODEL_VLM='your-vision-language-model'
# Optional model overrides for rfs rebuild-editable.
$env:RFS_REBUILD_LAYOUT_MODEL=$env:MODEL_VLM
$env:RFS_REBUILD_CONTROL_MODEL=$env:MODEL_VLM
$env:RFS_REBUILD_SEMANTIC_MODEL=$env:MODEL_VLM
$env:RFS_PROFESSIONAL_REBUILD_MODEL=$env:MODEL_VLM
# Required only for --asset-mode api slot-level image generation.
$env:GEMINI_API_KEY=$env:API_KEY
$env:GEMINI_GEN_IMG_URL='https://your-provider/v1beta/models/your-image-model:generateContent'Best quality, using VLM layout/control/semantic planning plus generated slot assets:
rfs rebuild-editable `
--reference "C:\path\figure.png" `
--out "output\editable_rebuild" `
--asset-mode api `
--layout-mode hybrid `
--control-mode hybrid `
--text-mode ocr `
--export-previewHigher-quality scripted mode, where the VLM first writes a controlled Figure DSL that mimics the best specialized rebuild scripts:
rfs rebuild-editable-pro `
--reference "C:\path\figure.png" `
--out "output\editable_rebuild_pro" `
--asset-mode api `
--repair-rounds 2 `
--export-previewThe pro workflow writes professional_rebuild_script.dsl.json; edit that file
and rerun with --compile-only to recompile without rerunning VLM planning or
image-generation API calls. Use --benchmark-out output\known_good_rebuild to
write professional_gap_report.json against a specialized-script output. Use
--repair-mode vlm only when you want the VLM to apply constrained DSL patches
after preview comparison.
Lower-cost mode, using VLM structure planning but keeping reference crops as slot assets:
rfs rebuild-editable `
--reference "C:\path\figure.png" `
--out "output\editable_rebuild_crop" `
--asset-mode crop `
--layout-mode hybrid `
--control-mode hybridOffline smoke test with no API:
rfs rebuild-editable `
--reference "C:\path\figure.png" `
--out "output\editable_rebuild_placeholder" `
--asset-mode placeholder `
--layout-mode heuristic `
--control-mode heuristicThe rebuild workflow writes reference_geometry_overlay.png and
reference_controls_overlay.png for inspection. If the automatic layout or
arrows need correction, edit reference_geometry.json or
reference_controls.json, then recompile without another API call:
rfs rebuild-editable `
--reference "C:\path\figure.png" `
--out "output\editable_rebuild" `
--compile-onlyTo validate whether VLM planning is improving a given reference image, run the
paired evaluator. It creates case_heuristic and case_vlm outputs under the
same directory and defaults to --asset-mode crop so it does not spend image
generation credits:
rfs rebuild-editable-eval `
--reference "C:\path\figure.png" `
--out "output\editable_rebuild_eval" `
--asset-mode crop `
--export-previewReview rebuild_vlm_eval_summary.json,
case_heuristic/reference_geometry_overlay.png, and
case_vlm/reference_geometry_overlay.png before running a full
--asset-mode api rebuild. Each rebuild also writes
rebuild_vlm_validation_report.json with layout/control/semantic validation
counts, fallback status, and API request counts.
Recommended real run:
rfs make-framework `
--paper "C:\path\paper.pdf" `
--reference "C:\path\reference.png" `
--out "output\paper_reference_image2" `
--slot-count 40 `
--slot-source reference-primary `
--complexity-profile reference-dense `
--candidates-per-slot 4 `
--locator-mode vlm `
--control-localizer-mode hybrid `
--arrow-style-mode reference `
--prompt-plan-mode vlm `
--prompt-plan-workers 8 `
--asset-mode image2 `
--asset-workers 6 `
--asset-retries 3 `
--asset-review-mode heuristic `
--critic-mode heuristic `
--text-extractor-mode ocr `
--ocr-engine paddle `
--ocr-lang en_ch `
--jsonUse lower worker counts if your API provider rate-limits requests.
input archive -> paper brief -> reference_geometry.json/reference_control_candidates.json ->
slot_overlay.png/reference_control_overlay.png -> reference_controls.json ->
arrow_style_profile.json/selected_arrow_routes.json/arrow_quality_report.json ->
reference_style_profile.json/style_sheet.md -> layout_plan.json -> figure_program.json ->
slot_visual_spec.json -> reference_slot_prompt_brief.json -> slot_prompt_plan.json ->
multi-candidate slot assets -> asset_quality_report.json -> asset_complexity_report.json ->
asset_visual_review.json/contact sheets -> editable_composition.pptx -> PDF/PNG export ->
visual_critic_iter_0.json -> critic_report.md -> validation
Optional Presentations-plugin QA can run after validation:
editable_composition.pptx -> rfs presentations-qa -> presentations_plugin_qa_report.json/.md
Optional whole-image Creator/Judge refinement can run before conversion:
structured Ground Truth -> Creator Agent candidates -> Online Judge repair feedback ->
Frozen Judge acceptance -> approved_image.png -> existing editable PPTX workflow
See docs/coevolution.md for the rfs coevolve-image command and Ground Truth contract.
The long-term data, training, evaluation, rollout, and Creator-coordination plan is maintained in docs/judge_model_training_roadmap.md.
User content hierarchy, emphasis, aesthetic, reference-image, and A/B preference criteria can be collected with docs/aesthetic_ground_truth_questionnaire.md.
Key rules:
- The reference image is the source of truth for layout, local visual object choice, color, visual rhythm, and arrow logic when
--slot-source reference-primaryis used. - The paper provides scientific terminology and concept mapping; it should not override the reference image into a generic template.
- Arrows, connector lines, dashed loops, panel frames, labels, formulas, and critical text are PPT editable objects, not image assets.
- Arrow/control localization is reference-driven: CV detects candidates, overlays label them, optional VLM binding assigns source/target semantics, and the PPT compiler renders editable connectors.
- Arrow styling is reference-preserving: it may soften line caps, assign bundle IDs, vary widths/dashes, and report aesthetics, but it must not replace reference-image flow logic with a generic router.
- Obstacle-aware routing is fallback-only: it may synthesize orthogonal paths for missing routes or
route_policy=fallback_reroute_allowed, but it must not rewrite reference-locked paths. - Aesthetic mode is experimental:
--arrow-style-mode aestheticmay offset reference-locked arrows only when the route explicitly opts in, only withinreference_tunnel_percent; it records the original path and must keepreference_tunnel_preserved=true. - Normal non-legend slots should be dense mini scientific scenes/cards with layered objects and micro-details, not simple centered icons.
- Generated images are inserted with no semantic cropping.
- The Presentations plugin is QA-only in this project. It can import/render a PPTX, extract layout JSON, and expose renderer/font/connector drift, but RFS remains the authoritative compiler for reference-locked geometry and connector-heavy figures.
Optional arrow-beautification pass:
rfs make-framework `
--paper "C:\path\paper.pdf" `
--reference "C:\path\reference.png" `
--out D:\ResearchFigureStudio\output\aesthetic_experiment `
--slot-count 40 `
--slot-source reference-primary `
--control-localizer-mode hybrid `
--arrow-style-mode aesthetic `
--prompt-plan-mode vlm `
--asset-mode image2For publication-safe comparison, keep both --arrow-style-mode reference and
--arrow-style-mode aesthetic outputs. Use the aesthetic version only when the
small reference-tunnel deviations improve readability without changing the
reference image's flow logic.
Optional Presentations-plugin QA for an existing output:
rfs presentations-qa `
--out "output\paper_reference_image2" `
--scale 2 `
--jsonOptional QA during a new run:
rfs make-framework `
--paper "C:\path\paper.pdf" `
--reference "C:\path\reference.png" `
--out "output\paper_reference_image2" `
--slot-source reference-primary `
--asset-mode image2 `
--presentations-qa `
--jsonIf the plugin reports autoRouteConnectorPx failed, treat that as a QA signal,
not as a reason to let the plugin rewrite the figure. The PPTX generated by RFS
remains the source of truth.
A valid image-rich framework run should include:
input_manifest.jsonpaper_brief.md/paper_brief.jsonreference_geometry.jsonreference_control_candidates.jsonslot_overlay.pngreference_control_overlay.pngreference_controls.jsonreference_text_geometry.jsontext_program.jsonocr_text_quality_report.jsontext_alignment_report.jsonarrow_style_profile.jsonselected_arrow_routes.jsonarrow_quality_report.jsonreference_style_profile.jsonstyle_sheet.mdlayout_plan.jsonfigure_program.jsonslot_visual_spec.jsonreference_slot_prompt_brief.jsonslot_prompt_plan.jsonprompts.mdreference_slot_crops/<slot_id>.pngassets/*.pngwith at least 25 selected non-arrow image assetsasset_candidates/*/candidate_*.pngasset_quality_report.jsonasset_complexity_report.jsonasset_visual_review.jsonasset_contact_sheet.pngasset_candidate_contact_sheet.pngeditable_composition.pptxreview.pdfandfinal_600dpi.pngwhen local export is availablevisual_critic_iter_0.jsonalignment_review.mdcritic_report.md
Optional QA outputs:
presentations_plugin_qa_report.jsonpresentations_plugin_qa_report.mdpresentations_plugin_qa_workspace/when using the default workspace
Run validation:
rfs validate --out "output\paper_reference_image2" --json
python codex-skills\research-figure-making\scripts\validate_framework_outputs.py "output\paper_reference_image2"This repository includes the Codex skill under:
codex-skills/research-figure-making
To install it locally into Codex:
$dst = Join-Path $env:USERPROFILE ".codex\skills\research-figure-making"
if (Test-Path $dst) { Remove-Item -Recurse -Force $dst }
Copy-Item -Recurse "codex-skills\research-figure-making" $dstThe skill documents the full research-figure workflow and includes the standalone framework-output validator.
python -m compileall -q rfs
python -m unittest discover -s tests -q
python -m py_compile codex-skills\research-figure-making\scripts\validate_framework_outputs.pyDo not commit:
output/- papers, manuscripts, private datasets, or user reference images
- generated PPTX/PDF/PNG/JPG/SVG assets
.envfiles or API keys- cache folders such as
__pycache__/or*.egg-info/
MIT License. See LICENSE.
