Repository navigation
feat: performance optimization phase 3 — zero-copy frame pool, per-worker SPSC - #25
Conversation
…rker SPSC, RSS affinity Multi-core pipeline redesign addressing the remaining performance gaps: - P3.1: FramePool slab allocator — pre-allocated contiguous buffer of N×frame_size bytes with lock-free SPSC free list. Eliminates all per-packet heap allocation on the RX→Worker hot path. - P3.2-P3.3: FrameRef (pool_idx + len) replaces Vec<u8> in worker SPSC rings. Frames are passed by 8-byte reference instead of cloned Vec, zero-copy after initial DMA. - P3.4: Per-worker SPSC app rings replace shared MPSC app_ring. recv_from() polls round-robin across worker rings, eliminating CAS contention that scaled poorly with worker count. - P3.6: RSS-aware worker affinity — direct queue-to-worker mapping preserves flow locality from NIC RSS hashing. - P3.8: Comprehensive unit tests for FramePool (alloc/free cycles, pool exhaustion, producer-consumer with SPSC ring, concurrent access). https://claude.ai/code/session_01AX56ktjxp4ybrKFqAg7mMu
Mark P3.1-P3.4, P3.6, P3.8 complete. Add preliminary Run #3 entry to perf-test-log.md with pending benchmark results (perf-tests.yml running). https://claude.ai/code/session_01AX56ktjxp4ybrKFqAg7mMu
[Perf] Stage: DeployDeploying |
[Perf] Stage: Instances Ready
|
[Perf] Stage: TRex ConfigStarting TRex configuration (MAC discovery + NIC binding)... |
[Perf] Stage: TRex Config OK
|
[Perf] Stage: TRex StartedTRex server running. Beginning benchmarks... |
[Perf] DUT ReadyDUT instance |
[Perf] Stage: Benchmark (1/5)Running |
[Perf] Benchmark Diag:
|
[CI] Stage: DeployInfrastructure ready.
|
[Perf] Benchmark Diag:
|
[Perf] Stage: Benchmark (2/5)Running |
[Perf] Benchmark Diag:
|
[Perf] Benchmark Diag:
|
[Perf] Stage: Benchmark (3/5)Running |
[Perf] Benchmark Diag:
|
[CI] Stage: SummaryAll tests PASSED. ARP seeding: kernel /proc/net/arp (automatic)
|
✅ Integration Tests Passed (Run 23035234544)Branch: Test Results
Application Logs (last 20 lines)receiver-echo-server.log sender-echo-server.log sender-test-client.log receiver-test-client-iperf.log sender-test-client-iperf.log Full Application Logs (last 200 lines each)receiver-echo-server.logsender-echo-server.logsender-test-client.logreceiver-test-client-iperf.logsender-test-client-iperf.log
|
[Perf] Benchmark Diag:
|
[Perf] Stage: Benchmark (4/5)Running |
[Perf] Benchmark Diag:
|
[Perf] Benchmark Diag:
|
[Perf] Stage: Benchmark (5/5)Running |
[Perf] Benchmark Diag:
|
[Perf] Benchmark Diag:
|
[Perf] Diag: testpmd logtestpmd output (last 30 lines) |
[Perf] Stage: Results[04:39:58] INFO Generating markdown summary... Performance Test Results — c5n.2xlargeCommit: 1400B packets
512B packets
64B packets
|
…rings Two critical fixes for Phase 3 performance: 1. FramePool::free() race condition: Changed from load+store to fetch_add for MPSC safety when multiple workers free concurrently. Increased free list to 2x capacity to prevent wraparound collisions. 2. True zero-copy: Workers now create AppPacket (FrameRef + payload offset) instead of ProcessedPacket (Vec<u8>). Frames stay in pool until recv_from() copies payload to user buffer and frees. Eliminates the per-packet heap allocation that negated FramePool's benefit. https://claude.ai/code/session_01AX56ktjxp4ybrKFqAg7mMu
Phase 3 Fix: True Zero-Copy PipelineThe initial Phase 3 benchmark showed a regression because:
Changes
Re-running perf tests: https://github.com/gspivey/dpdk-stdlib-rust/actions/runs/23036730290 |
[Perf] Stage: TRex Config OK
|
[Perf] Stage: TRex StartedTRex server running. Beginning benchmarks... |
[Perf] DUT ReadyDUT instance |
[Perf] Stage: Benchmark (1/5)Running |
[Perf] Benchmark Diag:
|
[Perf] Benchmark Diag:
|
[Perf] Stage: Benchmark (2/5)Running |
[Perf] Benchmark Diag:
|
[CI] Stage: SummaryAll tests PASSED. ARP seeding: kernel /proc/net/arp (automatic)
|
✅ Integration Tests Passed (Run 23077712886)Branch: Test Results
Application Logs (last 20 lines)receiver-echo-server.log sender-echo-server.log sender-test-client.log receiver-test-client-iperf.log sender-test-client-iperf.log Full Application Logs (last 200 lines each)receiver-echo-server.logsender-echo-server.logsender-test-client.logreceiver-test-client-iperf.logsender-test-client-iperf.log
|
[Perf] Benchmark Diag:
|
[Perf] Stage: Benchmark (3/5)Running |
[Perf] Benchmark Diag:
|
[Perf] Benchmark Diag:
|
[Perf] Stage: Benchmark (4/5)Running |
[Perf] Benchmark Diag:
|
[Perf] Benchmark Diag:
|
[Perf] Stage: Benchmark (5/5)Running |
[Perf] Benchmark Diag:
|
[Perf] Benchmark Diag:
|
[Perf] Diag: testpmd logtestpmd output (last 30 lines) |
[Perf] Stage: Results[02:47:41] INFO Generating markdown summary... Performance Test Results — c5n.2xlargeCommit: 1400B packets
512B packets
64B packets
|
…dmap Performance test results from GH Actions run 26204734648 show no regression from the IPv6 outer encap feature: - rust-dpdk 64B/700K: 699,000 RX (0.1% drop) — identical to Run #18 - rust-dpdk 512B/700K: 699,000 RX (0.1% drop) — identical to Run #18 - NIC instrumentation self-check: OK (zero drift) README roadmap updated to move 'Encap: IPv6 outer' from Planned to Done.
Append performance test results from GH Actions run 26227356354 (Graviton, TRex). No regressions detected — IPv6 fallback path adds zero overhead to IPv4 traffic. Mark IPv6 sub-task #2 (UDP over IPv6 checksum) as complete in the README roadmap.
Performance test results from GH Actions run 26633424088: - rust-dpdk 700K/64B: 695,587 RX (0.6% drop) — no regression vs Run #25 - rust-dpdk 700K/512B: 693,903 RX (0.9% drop) — no regression - tokio-dpdk caps at ~311K PPS — unchanged - native-dpdk baseline: 698,590 at 700K/64B Conclusion: IPv6 socket address support is performance-neutral.
## Roadmap Item `ROADMAP.md` item #25: `dpdk-stdlib-tcp`: Engine — on_command Spec: `.kiro/specs/tcp-support/` · tasks 5.18, 5.19 ## Changes Implements `TcpEngine::on_command` for the remaining three command variants: ### Shutdown (task 5.19) - **Shutdown::Write / Both**: Flush-before-FIN semantics — drains tx_ring into send_buf immediately, sets `fin_pending` flag. The `on_tick` TX drain path transmits all remaining data then emits FIN once send_buf is empty. - **Shutdown::Read**: Sets EOF on ConnectionHandle so app-side reads return 0. No FIN sent. - State transitions: Established → FinWait1, CloseWait → LastAck. ### Close (task 5.19) - **linger = None or linger > 0**: Graceful FIN (same as Shutdown::Write). App-side blocking handled via condvar. - **linger = Some(Duration::ZERO)**: Send RST immediately, discard all unsent data, remove TCB, latch ConnectionAborted. ### SetOption (task 5.19) - Nodelay, Keepalive (enable/disable clears timer), Linger (updates both TCB and handle), RecvBufSize, SendBufSize, ReuseAddr, Ttl. - ReadTimeout/WriteTimeout/Nonblocking are app-side only (no engine action needed). ### Infrastructure - `send_fin()` helper: builds FIN+ACK frame, advances snd_nxt, transitions state, arms RTO. - `fin_pending` handling in `on_tick`: after TX drain completes with no unsent data remaining, emit FIN via `send_fin()`. ## Tests Added 28 new unit tests covering: - Shutdown with empty/pending send_buf, Read/Write/Both variants - FIN sequence number consumption and RTO arming - Close with default/non-zero/zero linger - RST on linger=0 discards pending data - SetOption for all engine-relevant variants - Flush-before-FIN integration (data segment + FIN in correct order) - Edge cases: nonexistent key, shutdown in SYN_SENT ## Test Results All 192 dpdk-stdlib-tcp tests pass. Full workspace builds and all tests pass. --------- Co-authored-by: Agent Router <agent@agent-router.dev>
Summary
Multi-core pipeline redesign (Phase 3) targeting the remaining performance gaps at 350K-700K PPS:
FrameRef(pool_idx + len) replacesVec<u8>in worker SPSC rings. Frames passed by reference, not cloned.app_ringwith per-worker SPSC rings.recv_from()polls round-robin, eliminating CAS contention.Expected Impact (based on Phase 2 baseline)
Phase 2 Baseline (to beat)
Test plan
https://claude.ai/code/session_01AX56ktjxp4ybrKFqAg7mMu