Summary
Enabling --speedtest-url makes the check pipeline almost never commit a result: throughput drops from ~1.5 items/s to ~0.011 items/s (~130×) on a small list, and on a large run it sits at Checked 0/338707 for 35 minutes with zero observations written. Removing --speedtest-url and changing nothing else makes the same database check at 6–11 items/s.
Environment
- Proxy Workbench 3.0.3, installed from the release wheel
proxy_workbench-3.0.3-py3-none-any.whl (sha256 591c14592521b47d86800c035eca663201d3e69857e24881b8947c84c0eb9146, matches the published SHA256SUMS), into a clean venv.
- Windows 10 22H2 (19045), Python 3.11.3, 16 cores.
- Network note (possibly relevant): this uplink is behind a TUN that accepts every outbound TCP connect in ~5 ms and then blackholes the traffic. Proof: TCP-connect to 80 random candidates from the collected database: 80/80 "open", median 4.6 ms. So "connected" here means nothing and anything without a hard timeout will hang instead of failing fast. This makes the per-stage budgets matter a lot — but the contrast below is on the same machine and network, with the speedtest flag as the only variable.
Steps to reproduce
Collect first (this part is fine, 106 feeds):
proxy-workbench collect --data <DIR> --source-timeout 20
# -> 106 sources: 103 ok, 2 TimeoutError, 1 RemoteProtocolError
# -> 694500 raw rows -> 338707 unique candidates
A. Large run with a speedtest — stalls at zero
proxy-workbench scan --data <BIG_DIR> --url https://example.com/ --attempts 1 \
--timeout 8 --connect-timeout 3 --workers 256 --max-requests 20000 \
--speedtest-url "https://speed.cloudflare.com/__down?bytes=2000000" \
--speedtest-bytes 2000000
Observed (twice, two independent jobs, one on each of two consecutive nights of data):
| job |
wall time |
result line |
observations |
results |
job_item moved |
job-3799c320d82e8a84 |
2090 s (then terminated) |
Checked 0/338707; matching 0; saved 0 |
0 |
0 |
pending 338707 |
job-38251cf1c365efaa |
≥120 s |
Checked 0/338707; ... |
0 |
0 |
pending 338707 |
During the 2090 s run the worker held 34 established sockets and had accumulated ~22 s of user CPU, i.e. it was alive and connected but never committed a single item.
B. Same database, same flags, speedtest removed — works
proxy-workbench scan --data <SAME_BIG_DIR> --url https://example.com/ --attempts 1 \
--timeout 8 --connect-timeout 3 --workers 256 --max-requests 20000 --deadline 120
Observed: Checked 186/338707; 7.1 proxies/s after 50 s, 130 observations / 130 results written. A separate 5 000-candidate run behaved the same way (479 done in 45 s, ~11/s, worker RSS 59 MB).
C. Small list, no prefilter, speedtest on — reproduces in miniature
proxy-workbench run --no-sources --input 30-proxies.txt --data <SMALL_DIR> \
--url https://example.com/ --attempts 1 --timeout 8 --workers 30 \
--prefilter 0 --deadline 90 \
--speedtest-url "https://speed.cloudflare.com/__down?bytes=2000000" \
--speedtest-bytes 2000000
Observed: Checked 1/30; matching 1; saved 1 — 1 item committed in 90 s, the other 29 never did, although 30 workers were idle-ish and --timeout was 8 s. The single committed observation took 16.2 s (started_at → finished_at), and its stored speed payload is:
{"state": "error", "mbps": null, "bytes": 0, "chunks": 0, "ttfb_ms": null,
"transfer_ms": null, "total_ms": null, "code": "WHOLE_PROBE_TIMEOUT",
"detail": "весь срок 16 с исчерпан", "connection": "cold", "error": "WHOLE_PROBE_TIMEOUT"}
The identical command without --speedtest-url completed 30/30 in ~20 s (1.5–1.7 items/s) on the same machine.
Expected
Enabling the optional speed measurement should slow a run down and leave speed.mbps empty for proxies that cannot transfer bytes — not stop results from being committed. Running the three variants above on the same host, I'd expect all of them to make progress, with the speedtest one only slower proportionally to --speedtest-bytes.
Hypothesis
- The speed probe has its own budget (
WHOLE_PROBE_TIMEOUT, ~16 s per proxy here regardless of --timeout 8), which is fine by itself, but an item appears not to be committed until the speed stage finishes — so on a network where the transfer never completes, items pile up instead of being recorded as "checked, speed unavailable".
- At 338 707 candidates this shows up as a total stall (
0/338707), which is the scariest symptom: the progress line and job state say the job is running while the database receives nothing.
Workaround
Omit --speedtest-url / --speedtest-bytes; the rest of the pipeline then behaves as documented. The speed/bandwidth sort keys and the Mbit/s column in the UI are unavailable as a result.
Notes
- Both large jobs ended as
partial in the job table (the second one was terminated by me after ~35 min of zero progress; the first one likewise), but the measured fact is that in 2090 s not one of 338 707 items left pending.
- I'm happy to re-run any variation, or attach a
diagnose bundle, if that helps localise it. I did not include the database (≈500 MB) or personal paths here.
RU, кратко: при включённом --speedtest-url проверка практически перестаёт записывать результаты: на 30 прокси в базу попал 1 результат за 90 с против 30 за 20 с без флага, а на 338 707 кандидатах вывод 35 минут стоял на Checked 0/338707 при нуле наблюдений. Без --speedtest-url та же база проверяется на 6–11 адресов/с. Похоже, элемент не коммитится, пока не отработает этап замера скорости, а он в такой сети висит до WHOLE_PROBE_TIMEOUT (16 с).
Summary
Enabling
--speedtest-urlmakes the check pipeline almost never commit a result: throughput drops from ~1.5 items/s to ~0.011 items/s (~130×) on a small list, and on a large run it sits atChecked 0/338707for 35 minutes with zero observations written. Removing--speedtest-urland changing nothing else makes the same database check at 6–11 items/s.Environment
proxy_workbench-3.0.3-py3-none-any.whl(sha256591c14592521b47d86800c035eca663201d3e69857e24881b8947c84c0eb9146, matches the publishedSHA256SUMS), into a clean venv.Steps to reproduce
Collect first (this part is fine, 106 feeds):
A. Large run with a speedtest — stalls at zero
Observed (twice, two independent jobs, one on each of two consecutive nights of data):
observationsresultsjob_itemmovedjob-3799c320d82e8a84Checked 0/338707; matching 0; saved 0pending338707job-38251cf1c365efaaChecked 0/338707; ...pending338707During the 2090 s run the worker held 34 established sockets and had accumulated ~22 s of user CPU, i.e. it was alive and connected but never committed a single item.
B. Same database, same flags, speedtest removed — works
Observed:
Checked 186/338707; 7.1 proxies/safter 50 s, 130 observations / 130 results written. A separate 5 000-candidate run behaved the same way (479 done in 45 s, ~11/s, worker RSS 59 MB).C. Small list, no prefilter, speedtest on — reproduces in miniature
Observed:
Checked 1/30; matching 1; saved 1— 1 item committed in 90 s, the other 29 never did, although 30 workers were idle-ish and--timeoutwas 8 s. The single committed observation took 16.2 s (started_at→finished_at), and its storedspeedpayload is:{"state": "error", "mbps": null, "bytes": 0, "chunks": 0, "ttfb_ms": null, "transfer_ms": null, "total_ms": null, "code": "WHOLE_PROBE_TIMEOUT", "detail": "весь срок 16 с исчерпан", "connection": "cold", "error": "WHOLE_PROBE_TIMEOUT"}The identical command without
--speedtest-urlcompleted30/30in ~20 s (1.5–1.7 items/s) on the same machine.Expected
Enabling the optional speed measurement should slow a run down and leave
speed.mbpsempty for proxies that cannot transfer bytes — not stop results from being committed. Running the three variants above on the same host, I'd expect all of them to make progress, with the speedtest one only slower proportionally to--speedtest-bytes.Hypothesis
WHOLE_PROBE_TIMEOUT, ~16 s per proxy here regardless of--timeout 8), which is fine by itself, but an item appears not to be committed until the speed stage finishes — so on a network where the transfer never completes, items pile up instead of being recorded as "checked, speed unavailable".0/338707), which is the scariest symptom: the progress line andjobstate say the job is running while the database receives nothing.Workaround
Omit
--speedtest-url/--speedtest-bytes; the rest of the pipeline then behaves as documented. Thespeed/bandwidthsort keys and the Mbit/s column in the UI are unavailable as a result.Notes
partialin thejobtable (the second one was terminated by me after ~35 min of zero progress; the first one likewise), but the measured fact is that in 2090 s not one of 338 707 items leftpending.diagnosebundle, if that helps localise it. I did not include the database (≈500 MB) or personal paths here.RU, кратко: при включённом
--speedtest-urlпроверка практически перестаёт записывать результаты: на 30 прокси в базу попал 1 результат за 90 с против 30 за 20 с без флага, а на 338 707 кандидатах вывод 35 минут стоял наChecked 0/338707при нуле наблюдений. Без--speedtest-urlта же база проверяется на 6–11 адресов/с. Похоже, элемент не коммитится, пока не отработает этап замера скорости, а он в такой сети висит доWHOLE_PROBE_TIMEOUT(16 с).