Conversation
davidism
force-pushed
the
automatic-options
branch
from
May 2, 2026 04:06
5565391 to
80669a0
Compare
yaoyi0427
added a commit
to yaoyi0427/flask
that referenced
this pull request
May 5, 2026
Document the provide_automatic_options parameter and other arguments. This helps users understand how URL routing works in Flask. Refs: pallets#5918 Co-authored-by: 姚毅 <yaoyi@yaoyideMac-mini.local>
davidism
force-pushed
the
automatic-options
branch
from
August 11, 2026 20:21
80669a0 to
a82e942
Compare
Dantes19xx
added a commit
to Dantes19xx/oss_copilot
that referenced
this pull request
Sep 8, 2026
Linear graph: fetch_diff (via in-process MCP client -> own MCP server) -> analyze_diff (gpt-4o-mini, structured output) -> aggregate_review. No branching/loops/human-in-the-loop yet (stage 5). Verified end-to-end against a real PR (pallets/flask#5918).
Dantes19xx
added a commit
to Dantes19xx/oss_copilot
that referenced
this pull request
Sep 8, 2026
New backend/graph/vision.py: extracts image URLs from a PR body (markdown, <img>, GitHub asset hosts) and analyzes each via gpt-4o-mini vision. Wired into the review branch as a genuine conditional edge after fetch_diff (route_after_fetch_diff: "vision" vs "context") -> new analyze_screenshots node -> folded into aggregate_review's summary. Verified live: a real merged PR with a demo GIF in its description (facebook/docusaurus#7036) takes the vision path and the analysis shows up in the final review; a PR with no images (pallets/flask#5918) still takes the short "context" path with no regression.
Dantes19xx
added a commit
to Dantes19xx/oss_copilot
that referenced
this pull request
Sep 11, 2026
backend/graph/cache.py: FileCache, a small on-disk JSON cache keyed by sha256 of canonicalized inputs. Deliberately exact-match, not semantic/similarity-based — for code review, two 95%-similar diffs can still differ in the one line with the actual bug, so a similarity cache risks confidently serving the wrong review. Wired into two genuinely repeated calls: - backend/rag/embeddings.py: embed_texts() now caches per-text within a batch (only the real API call is @Traceable, so a cache hit correctly produces no LangSmith trace). Embeddings are a pure deterministic function of (model, text), so caching is unconditional. - backend/graph/nodes_review.py: caching lives in the analyze_file graph node only, not inside review_file() itself — eval harnesses (run_evals/run_ab/run_temperature) call review_file() directly and must always get a fresh model call, or the cache would silently mask the run-to-run variance already documented as a finding in stages 9 and 11. Cache key includes a hash of the current system prompt, so editing the prompt auto-invalidates old entries. Verified live: running the same real PR (pallets/flask#5918) twice went from 7 OpenAI calls (cold) to 1 (warm) — only intake classification isn't cached, by design.
Dantes19xx
added a commit
to Dantes19xx/oss_copilot
that referenced
this pull request
Sep 19, 2026
FileReviewOutput.comments is now list[ReviewComment] (code_line + text)
instead of list[str]. aggregate_review shows each reviewed file's actual
diff (added/removed lines) followed by its comments, each prefixed with
the resolved line number ("L12:") or "General" for file-wide remarks.
Two real bugs found by testing against live model calls, not assumed:
- Asking the model to compute the line number itself from the diff's hunk
header was off by 1-2 lines even on a single simple pure-addition hunk
(e.g. flagged eval(data) as line 4 when it was actually line 6).
Switched to deterministic resolution instead: the model quotes the exact
source line (code_line), and _new_file_line_map()/_resolve_comment_line()
parse the diff themselves to find the real line number — arithmetic an
LLM does unreliably, done in code instead.
- The model also doesn't reliably preserve a quoted line's leading
whitespace despite being told to, which silently broke exact-match
resolution. Added a whitespace-stripped fallback map, tried only after
the exact one so two lines sharing stripped content still resolve to the
right one when the model does quote precisely.
Updated run_evals.py/run_ab.py for the new comment shape (judge() reads
.text, JSON output uses .model_dump()) and reran the full 30-example
golden dataset, since the prompt genuinely changed (auto-invalidated the
review cache via its prompt-version hash). Numbers moved (0.83->0.80
accuracy, 4.13->3.73 judge) — documented honestly in EVALS.md as a 4th
data point in the existing run-to-run-variance discussion, including a
control test (same 8 clean examples, full vs. shortened anchoring
instruction, same 6/8 false-positive count either way) that rules out
prompt length as the driver without claiming the drop is proven noise.
Verified end-to-end on a real PR (pallets/flask#5918, 13 real comments
across 5 files) with line numbers spot-checked against the actual diff,
in both languages, and confirmed the now much-longer draft still scrolls
cleanly inside the frontend's card instead of overflowing the page.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Implement
provide_automatic_optionsas a regular part of the routing and view dispatch, rather than as a special check on every request.Can be updated to do
except DuplicateRuleErrorinstead ofExceptiononce Werkzeug 3.2 is out. Relies on the duplicate check to not add the internal route multiple times, for example if@app.getand@app.postregister different view functions for the same URL. But nothing breaks if the rule is added multiple times, it's just messy.Adjusts
flask routesto no show the_automatic_optionsendpoint, as that would be really noisy.Also moves the static view out instead of using a lambda and weakref, to match how this new view is handled.
expands on #5917, fixes #5916