Repository navigation
Websocket incoming frame parsing fixes - #473
willmmiles wants to merge 6 commits into
Conversation
|
How do you folks feel about the state machine parser? Is the robustness worth the performance cost, or should we instead try buffering the header (spends more RAM and/or we're playing I don't have a good way to quantify the performance either, all of our benchmarks so far have focused on outgoing frames. |
That's a tricky question! I think we could test that with websocat. |
9ba38e6 to
7e0ee46
Compare
|
I'm working on some test cases for the fixes. Claude and I have found that the last fix (close reason handling) isn't really sufficient - we will have to defragment control frames for standards compliant operation. Stand by for more code. |
7e0ee46 to
6cca924
Compare
|
I read recently that modern AI is really, really good at finding all the bugs you ask it for. This is definitely turning out to be my experience here! I'm trying to validate the close-on-error semantics and it's turning in to a rabbit hole. There's a pernicious corner case with ... and there's more: in most cases where we close as a result of the client asking (like AsyncWebSocket), we must ack the bytes read before closing. Otherwise the TCP stack generates a RST-close instead of a FIN-close, as is required by the TCP protocol, to indicate that the remote client did not in fact consume all bytes. Even if it was safe to call I'm going to think about this one a bit -- wanted to share where it's at, though. |
it's like we need a "deferred" close ? |
Yup, that's one approach. Basically the solution space breaks down to:
Lots to think about. |
Use a state machine to process headers byte-by-byte so we can handle partial reception at any point.
If a control frame spans multiple TCP packets, buffer the data so that the frame can be processed once fully received. This ensures that the frame can be correctly handled instead of generating invalid PONG responses or overrunning the buffer with a disconnect reason.
6cca924 to
ddf5ce9
Compare
ESP32Async/AsyncTCP#124 implents this solution - it adds that guarantee that it's safe to I've pushed one more commit here that reorders things so it's "as safe as possible" with older AsyncTCP. This is as good as it'll get, I think. |
|
Reporting back with the real-Safari disconnect testing you flagged as untestable — done against Three disconnect paths from Safari (macOS, same machine the WS server sees):
Across the whole session: zero panics, zero reboots, zero error lines, no dropped-frame accounting anomalies, heartbeat cadence continuous. Heap floor during the Safari load + reconnect burst was ~7.9 KB with all responses completing. Honest caveat: this is black-box validation — I did not capture the wire bytes, so I cannot prove the specific "missing mask on close frame" Safari quirk actually fired; what I can say is that repeated real-Safari disconnects (all three styles, several rounds) never upset the parser. If you want wire-level evidence of that exact frame shape, I can arrange a capture. Combined with the earlier torn-frame injection suite (64 injections across header/mask/payload/close tears + fuzz, all clean, connection survives torn headers and keeps answering — reported in #481), our side has nothing blocking this PR. |
Thanks, that's still very helpful. As part of the patch development process my harness and I ended up building a byte-by-byte socket test sequence for validating that the code works as designed, but I don't have any Apple devices on hand to see what any particular version of Safari actually sends. Thanks for giving it a try and providing feedback. |
@willmmiles : is it better to first complete, review and release the AsyncTCP PR ? |
There was a problem hiding this comment.
I did some manual tests and reviewed (helped with AI for spec correctness). I create 3 little comments below.
Thanks 👍
✅ RFC 6455 compliance analysis
| Spec reference | Requirement | PR behavior |
|---|---|---|
| §5.2 | Base framing: FIN, opcode, mask bit, 7/16/64-bit payload length | State machine parses all header forms byte-at-a-time; length assembly correct for all three forms |
| §5.4 | Fragmentation: FIN + continuation opcode 0 | message_opcode preserved across fragments; final/num tracking correct |
| §5.5 | Control frames MUST be ≤ 125 bytes | New validation in Length state → close with 1002 ✅ |
| §5.5 | Reserved control opcodes 0xB–0xF | Unknown control opcodes → close with 1002 ✅ |
| §5.5.1 | Close frame: 2-byte status code (network order) + UTF-8 reason | data[0]<<8 + data[1] ✅; strnlen(reason, datalen-2) fixes old strlen over-read ✅ |
| §5.5.1 | Server MUST echo close frame in response | _queueControl(WS_DISCONNECT, data, datalen) echoes received payload ✅ |
| §5.5.1 | If both sides sent close → close TCP connection | _status == WS_DISCONNECTING → _client->close() ✅ |
| §5.5.2 | PONG must echo PING payload | _queueControl(WS_PONG, data, datalen) ✅ |
| §5.4 | Control frames MAY be interleaved in fragmented messages | Interleaved control frames don't clobber message_opcode ✅ |
| §7.4.1 | Status codes 1000–1011 | New AwsCloseCode enum matches; server uses 1002 (protocol error) and 1011 (internal error) correctly ✅ |
| §7.1.7 | Fail connection on protocol violation | All violations close with 1002 ✅ |
⚠️ Not validated (pre-existing gaps, same as old code — optional follow-up hardening)
| Spec reference | Requirement | Status |
|---|---|---|
| §5.2 | RSV1–3 bits MUST be 0 (no extensions negotiated) | Not checked — lenient |
| §5.2 | Reserved data opcodes 0x3–0x7 | Not checked — silently treated as data |
| §5.5 | Control frames MUST NOT be fragmented (FIN=1) | Not checked |
| §5.2 | Minimal-length encoding (e.g. 16-bit form for len < 126) | Not checked |
| §5.1 | Client→server frames MUST be masked | Not enforced — server accepts unmasked frames (lenient, common for embedded servers) |
| §7.4.1 | 1005/1006 MUST NOT appear on wire | Received codes not validated (passed to app as-is) |
| _pinfo.message_opcode = _pinfo.opcode; | ||
| } | ||
| // init frame number to 0 if only 1 frame or if this is the first frame of a fragmented message | ||
| if (_pinfo.final || datalen < _pinfo.len) { |
There was a problem hiding this comment.
| if (_pinfo.final || datalen < _pinfo.len) { | |
| // note: only for data frames; an interleaved control frame (ping/pong/close) must not reset the | |
| // fragment counter of an in-progress fragmented message | |
| if (_pinfo.opcode < WS_DISCONNECT && (_pinfo.final || datalen < _pinfo.len)) { |
_pinfo.num is the frame counter within a fragmented message, exposed to applications via AwsFrameInfo in WS_EVT_DATA events. The examples use it to distinguish "MSG START" (num == 0) from subsequent frames (examples/arduino/WebSocket/WebSocket.ino:195).
The problem: this block runs for every frame whose index == 0, including control frames (PING/PONG/close). Control frames always have final == true. So if a PING arrives between fragment 1 and fragment 2 of a fragmented text message, app sees num=0 on the continuation frame and may interpret it as a new message start.
Note: this is a pre-existing bug.
There was a problem hiding this comment.
Hm, I'm not sure this is quite right either - I think the code is confusing a fragmented message with a fragmented frame. Clearing based on datalen is probably also wrong because a frame fragmented between two packets in the middle will reset the counter for a message.
I'll have to think about this and come back to it.
There was a problem hiding this comment.
The fix is really about fragmented messages only (not paquets) so when num increases and a control frame arrives in between.
This is an edge case because WS message frames can be huge in size so on a MCU we barely can have the use case of a message split into frames: it could happen though for a sort of streaming application i guess ?
| } | ||
| } | ||
| if (_status == WS_DISCONNECTING) { | ||
| _status = WS_DISCONNECTED; |
There was a problem hiding this comment.
| _status = WS_DISCONNECTED; | |
| _status = WS_DISCONNECTED; | |
| _pstate = AwsParseState::Error; // terminal state: ignore any further data before the disconnect completes |
Parser enters terminal state; any trailing TCP data before disconnect completes is silently ignored instead of re-firing WS_EVT_ERROR
| return; // our object is now destroyed, so we must return immediately to avoid accessing any member | ||
| return false; // our object is now destroyed, so we must return immediately to avoid accessing any member | ||
| } else { | ||
| _status = WS_DISCONNECTING; |
There was a problem hiding this comment.
| _status = WS_DISCONNECTING; | |
| _status = WS_DISCONNECTING; | |
| _pstate = AwsParseState::Error; // terminal state: ignore any further data before the disconnect completes |
Parser enters terminal state; any trailing TCP data before disconnect completes is silently ignored instead of re-firing WS_EVT_ERROR
|
Hi @willmmiles! |
Thanks! I'll get to it a soon as I can - sorry I'm a bit spread thin at the moment. |
No worry! |
Fix handling of incoming frame headers that span multiple packets, a corner case processing very large frames, and an edge case handling disconnect frames that are split between packets.
Filed as draft for these open concerns: