Skip to content

Fix DTLS 1.2 stall and resend storm when a flight's records arrive out of order - #168

Closed
z3thon wants to merge 2 commits into
algesten:mainfrom
Alakazam-211:fix-dtls-reorder-resend-storm
Closed

z3thon wants to merge 2 commits into
algesten:mainfrom
Alakazam-211:fix-dtls-reorder-resend-storm

Conversation

@z3thon

@z3thon z3thon commented Oct 2, 2026 •

Copy link
Copy Markdown

Summary

When the records of one DTLS 1.2 datagram arrive out of order (for example, the client's final flight with Finished before ChangeCipherSpec), the receiving endpoint can stall: it only looks for the next handshake message at the head of its receive queue, and it only accepts a ChangeCipherSpec when that is the first unprocessed record. The handshake then cannot finish. Every retransmission of the peer's flight carries a duplicate of a message the endpoint already processed, and each such duplicate makes it resend its own whole flight at once. The peer treats that resend the same way, so both sides bounce flights back and forth at network speed, with no backoff, until the handshake deadline. This PR buffers out-of-order handshake messages by message_seq (RFC 6347 §4.2.2) and limits duplicate-triggered resends to one per retransmission timer period (RFC 6347 §4.2.4). Each fix has a regression test.

Impact

  • One reordered datagram is enough. Reordering within a single datagram's records is unusual. Reordering by middleboxes, multipath links, or a peer that packs records differently is ordinary.
  • Once stalled, the two endpoints exchange flights as fast as the network round trip allows. Between two dimpl endpoints over loopback we measured about 20,000 datagrams per second in each direction (300,000 to 535,000 datagrams in 15 s), until the handshake deadline expired. The test below shows the same pattern in a simulated network with 10 ms steps: about 100 datagrams per side in 3 s, against 2 to 5 for a normal handshake.
  • The handshake never completes, so the connection fails after the full handshake timeout.
  • The per-duplicate resend on its own (without the stall) also lets any burst of duplicated datagrams trigger one full flight per duplicate.

Reproduction

  1. Clone dimpl and check out this PR's branch:

    git clone https://github.com/algesten/dimpl
    cd dimpl
    git fetch origin pull/168/head:dtls-reorder-fix
    git checkout dtls-reorder-fix
  2. Put back the unfixed sources from main, keeping the new tests:

    git checkout origin/main -- src
  3. Run the two new tests:

    cargo test --test dtls12 -- reversed_final_flight duplicate_flights_resend --nocapture

    Both fail on the unfixed code:

    running 2 tests
    thread 'retransmit::dtls12_duplicate_flights_resend_at_most_once_per_timer_period' panicked at tests/dtls12/retransmit.rs:1075:5:
    assertion `left == right` failed: 20 duplicates should resend flight 4 (1 datagrams) once, server sent 20
      left: 20
     right: 1
    test retransmit::dtls12_duplicate_flights_resend_at_most_once_per_timer_period ... FAILED
    mtu 1150 cookie false duplicate false: client sent 99, server sent 97 datagrams, connected after None
    thread 'reorder::dtls12_reversed_final_flight_completes_without_resend_storm' panicked at tests/dtls12/reorder.rs:510:5:
    mtu 1150 cookie false duplicate false: resend storm: client sent 99, server sent 97 datagrams in 3 s (at most 30 each)
    test reorder::dtls12_reversed_final_flight_completes_without_resend_storm ... FAILED
    
    test result: FAILED. 0 passed; 2 failed; 0 ignored; 0 measured; 73 filtered out
    

    The handshake never completes (connected after None). In the 3 simulated seconds, each side sends about 100 datagrams. They start when the client's 1 s retransmission timer fires, and then go one per 10 ms simulation step: one per round trip. On a real network that is one per round trip, until the handshake deadline.

  4. Restore the fix and run the same command again:

    git checkout HEAD -- src
    cargo test --test dtls12 -- reversed_final_flight duplicate_flights_resend --nocapture

    Both pass:

    running 2 tests
    mtu 1150 cookie false duplicate false: client sent 2, server sent 2 datagrams, connected after Some(40ms)
    mtu 1150 cookie false duplicate true: client sent 2, server sent 2 datagrams, connected after Some(40ms)
    mtu 1150 cookie true duplicate false: client sent 3, server sent 3 datagrams, connected after Some(60ms)
    mtu 1150 cookie true duplicate true: client sent 3, server sent 3 datagrams, connected after Some(60ms)
    mtu 300 cookie true duplicate false: client sent 5, server sent 5 datagrams, connected after Some(60ms)
    mtu 300 cookie true duplicate true: client sent 5, server sent 5 datagrams, connected after Some(60ms)
    test retransmit::dtls12_duplicate_flights_resend_at_most_once_per_timer_period ... ok
    test reorder::dtls12_reversed_final_flight_completes_without_resend_storm ... ok
    
    test result: ok. 2 passed; 0 failed; 0 ignored; 0 measured; 73 filtered out
    

    The full suite (cargo test), cargo clippy --all-targets -- -D warnings and cargo fmt --all -- --check pass on the branch.

The first test, dtls12_reversed_final_flight_completes_without_resend_storm, runs two dimpl endpoints over an in-memory channel. It reverses the records of the client's first datagram that carries ChangeCipherSpec (so Finished arrives first, and the handshake messages arrive in reverse order). It runs with and without duplicating every datagram, and with three configurations (no cookie, cookie, MTU 300). It asserts that the handshake completes before the first retransmission timer, and that each side sends at most 30 datagrams in 3 s. The second test, dtls12_duplicate_flights_resend_at_most_once_per_timer_period, sends bursts of 20 duplicate flights to a waiting server and a waiting client. It asserts that each burst triggers one resend per timer period.

Root cause

All line numbers are on main (98f893d). The same code is at the lines in brackets in 0.7.4.

What happens, step by step (from the server's debug log in the test above):

  1. The reversed copy of the client's final flight arrives as one datagram: Finished (epoch 1), ChangeCipherSpec, CertificateVerify, ClientKeyExchange, Certificate. Engine::has_complete_handshake_with_seq (src/dtls12/engine.rs:622 [0.7.4: 604]) only looks at the first unhandled handshake in the queue. It returns false unless that handshake has the expected message_seq (:640 [622]). Here the first one is CertificateVerify and the server expects Certificate, so nothing is processed. (Engine::next_record, :706 [688], has the same head-of-queue rule for ChangeCipherSpec. It takes only the first unhandled record (:711 [693]), which here is the epoch 1 Finished that cannot be decrypted yet.)
  2. About 1 s later, the client's timer resends the final flight in order. The server processes Certificate and ClientKeyExchange from that copy. For CertificateVerify, has_complete_handshake_with_seq finds the new copy and returns true. Then next_handshake (:673 [655]) passes Handshake::defragment (src/dtls12/message/handshake.rs:150, same in both) an iterator over every later handshake in the queue. The records between the two copies (ChangeCipherSpec and the undecrypted Finished) carry no parsed handshakes, so the stale CertificateVerify from the reversed copy comes next in that iterator. defragment appends every following handshake with the same type and message_seq (:170). It concatenates both copies and fails with "Fragment length mismatch" (:189). Nothing is marked handled, so every later attempt fails the same way. The server is stuck waiting for CertificateVerify.
  3. Every later copy of the client's final flight carries a duplicate ClientKeyExchange. Engine::insert_incoming_handshake resends the server's whole flight at once for every such datagram (src/dtls12/engine.rs:286 [283]). That resend carries a duplicate ServerHelloDone, which makes the client resend its final flight at once, and so on. The two endpoints bounce flights at round-trip speed until the handshake deadline. DTLS 1.3 has the same per-duplicate resend (src/dtls13/engine.rs:393 [390]).

The fix

  • RFC 6347 §4.2.2 buffering (DTLS 1.2). The next handshake message is looked up by its message_seq anywhere in the receive queue (complete_handshake_fragments). Its fragments are collected in offset order, and duplicate fragments are skipped. defragment gets exactly the fragments that make up the message, never a second copy. Later messages stay queued until it is their turn. Records of an epoch whose keys are not in place yet are skipped, both for this lookup and for next_record, so an early Finished waits for the ChangeCipherSpec instead of blocking it. Queued copies of messages that were already processed are marked handled (discard_stale_handshakes), so they neither block the next message nor pin their datagram in the receive queue.
  • RFC 6347 §4.2.4 timer-driven retransmission (DTLS 1.2 and 1.3). A duplicate of the peer's previous flight may trigger at most one early resend per timer period. The budget is restored when a new flight begins or when the flight timer fires; the timer keeps its exponential backoff. Once the timers are stopped after the final flight, every retransmission of the peer's last flight is still answered, as §4.2.4 requires; the peer's own timer paces those.
  • The existing test suite passes unchanged.

sew-build added 2 commits October 1, 2026 11:27
…rrive out of order

A DTLS 1.2 endpoint looked for the next handshake message only at the head of its receive queue, and took a ChangeCipherSpec only when it was the first unprocessed record. When the records of one datagram arrive out of order (for example the client's final flight with Finished before ChangeCipherSpec and the handshake messages reversed), the head of the queue held a record that could not be processed yet: an epoch 1 Finished waiting for its keys, or a later handshake message. The handshake stalled, and every retransmission of the peer's flight then triggered an immediate resend of our own flight.

Look up the next handshake message by its message_seq anywhere in the receive queue and keep later messages buffered (RFC 6347 §4.2.2). Let records of the next epoch wait for the ChangeCipherSpec without blocking it, and mark queued copies of messages that were already processed as handled so they neither block the next message nor pin their datagram in the queue.

Add a regression test that reverses the records of the client's final flight, with and without duplicated datagrams, and bounds the datagrams each side sends.
Every datagram carrying a duplicate of the peer's previous flight (a ClientKeyExchange or ServerHelloDone in DTLS 1.2, a ClientHello or Finished in DTLS 1.3) made the endpoint resend its whole current flight at once. Two endpoints that both wait, for example on a stalled handshake, then bounce their flights back and forth at network speed until the handshake deadline, with no backoff.

Retransmission stays driven by the flight timer with exponential backoff (RFC 6347 §4.2.4, RFC 9147 §5.8.1). A duplicate may trigger one early resend per timer period; the budget is restored when a new flight begins or the timer fires. Once the timers are stopped after the final flight, each retransmission of the peer's last flight is still answered, as RFC 6347 §4.2.4 requires; the peer's own timer paces those.

Add a regression test that delivers bursts of duplicate flights to a waiting server and a waiting client.
@algesten

algesten commented Oct 2, 2026

Copy link
Copy Markdown
Owner

Hi @z3thon! Thanks for this.

DTLS 1.2 is flight oriented so, resends are supposed to retransmit entire flights. What I believe you found is a state poisoning bug + a resend flood scenario, that needs fixing, but not exactly like in this PR.

A UDP packet can contain multiple DTLS messages that are numbered by message_seq. Whilst dimpl should handle fragmentation, we don't have the ambition to handle reordering. I.e.

If a datagram with message_seq: [2,3,4] somehow is rewritten [3,4,2] - we will reject it. However we do want to handle [3,4], [2] as separate datagrams.

What you found is that [3,4,2] actually poisons the state making it impossible to recover and that duplicate fragments unconditional immediate resends. That needs fixing.

@algesten

algesten commented Oct 2, 2026

Copy link
Copy Markdown
Owner

Alternative take in #169

@algesten algesten closed this Oct 3, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants