Repository navigation
feat: add TRex-based performance testing framework - #17
Merged
Merged
Conversation
Add a complete performance benchmarking pipeline that measures UDP echo throughput, latency, and packet loss across 4 configurations: - rust-dpdk: Our Rust+DPDK stack (the main thing we're measuring) - native-dpdk: testpmd macswap (DPDK C baseline / ceiling) - rust-stdlib: Our echo app falling back to kernel UDP - plain-rust: Minimal std::net echo server (kernel baseline) Infrastructure: - PerfTestStack CDK (TRex c5n.2xlarge + DUT c5n.2xlarge, dual ENIs) - TRex AMI Packer template + build workflow - Manual-trigger perf-tests.yml workflow with configurable params Orchestration (scripts/run-perf-tests.sh): - Deploys stack, waits for SSM, configures TRex - Runs each DUT config sequentially (rebinds NIC between DPDK/kernel) - Collects environment info, networking diagnostics, app logs - Posts staged PR comments at each phase (same pattern as integration tests) - Aggregates JSON results + generates markdown summary table - Safety-net teardown on failure New app: apps/plain-echo — minimal std::net UDP echo (~30 lines, no deps beyond clap) for the kernel baseline measurement. https://claude.ai/code/session_01KQw8H7AZWpb5jQLSNpPiGr
[CI] Stage: DeployInfrastructure ready.
|
[CI] Stage: SummaryAll tests PASSED. ARP seeding: kernel /proc/net/arp (automatic)
|
✅ Integration Tests Passed (Run 22811666900)Branch: Test Results
Application Logsreceiver-echo-server.logsender-echo-server.logsender-test-client.logreceiver-test-client-iperf.logsender-test-client-iperf.log
|
…:read) gh run watch requires the checks:read permission which fine-grained PATs cannot grant. Replace with gh run view --json polling loop in ci-validate.sh and document the limitation in CLAUDE.md so future sessions avoid it. https://claude.ai/code/session_01KQw8H7AZWpb5jQLSNpPiGr
[CI] Stage: DeployInfrastructure ready.
|
[CI] Stage: SummaryAll tests PASSED. ARP seeding: kernel /proc/net/arp (automatic)
|
✅ Integration Tests Passed (Run 22812424171)Branch: Test Results
Application Logsreceiver-echo-server.logsender-echo-server.logsender-test-client.logreceiver-test-client-iperf.logsender-test-client-iperf.log
|
gspivey
pushed a commit
that referenced
this pull request
Apr 13, 2026
Records hardware benchmark results from the spawn_blocking elimination. tokio-dpdk improved from ~40K pps (Run #9) to 343K pps at 700K/64B (8.6x). Matches sync DPDK at ≤140K pps; remaining gap at >350K pps is from Tokio yield_now() scheduling latency, not framework overhead. https://claude.ai/code/session_01EpC1AocdV7TkvTiMXmtCkm
gspivey
pushed a commit
that referenced
this pull request
Apr 13, 2026
The spin-poll optimization (commit a02b721) did not close the high-rate tokio-dpdk performance gap. Hardware perf tests showed essentially flat results vs Run #17, and the NIC counters (imissed=0) disprove the original hypothesis that yield_now() scheduler latency was causing ring overflow. The real bottleneck is per-packet CPU cost in the async echo path (~1.4 μs overhead per packet vs sync DPDK), likely from 2× Mutex lock/unlock per packet and async_trait vtable dispatch. Spin-polling doesn't help because the empty-poll idle path is not what's limiting throughput. Reverts the spin-poll change in DpdkUdpSocket::recv_from and the matching benchmark code. Adds Run #18 to docs/perf-test-log.md documenting the negative result and future directions that could actually help (lock-free Arc<UdpSocket>, batch APIs, dedicated poll thread + mpsc). https://claude.ai/code/session_01EpC1AocdV7TkvTiMXmtCkm
gspivey
pushed a commit
that referenced
this pull request
Apr 17, 2026
No regressions detected. rust-dpdk matches or slightly exceeds Run #17 at all packet sizes and rates. The GUE Option::is_some() check in send/recv hot paths adds zero measurable overhead when GUE is not configured. https://claude.ai/code/session_01NmhvqXvaxmA3fSXRGrzPpb
gspivey
added a commit
that referenced
this pull request
Apr 17, 2026
* feat: GUE tunnel endpoint (Generic UDP Encapsulation) Implement RFC 8470-style L3-over-UDP encapsulation as a transparent tunnel mode on UdpSocket. When GUE is configured, the application's send_to/recv_from calls are automatically encapsulated/decapsulated through the tunnel. Wire format: [Outer Eth][Outer IPv4][Outer UDP:6080][GUE 4B][Inner IPv4][Inner UDP][Payload] Key changes: - New dpdk-udp/src/gue.rs: GUE header build/parse, frame encap/decap functions - UdpSocket.set_gue()/gue(): per-socket tunnel configuration - TX path: builds GUE-encapsulated frames, ARP resolves tunnel remote endpoint - RX path: decapsulates matching GUE frames, returns inner src address to app - NetworkConfig.with_gue(): builder integration parallel to VLAN - MTU check accounts for 32-byte GUE overhead (outer IP + outer UDP + GUE header) - 23 unit tests: header codec, frame roundtrip, socket-level decap, PPS benchmark - README roadmap updated: GUE moved from Planned to Done https://claude.ai/code/session_01NmhvqXvaxmA3fSXRGrzPpb * docs: add Run #19 perf results — GUE tunnel endpoint regression check No regressions detected. rust-dpdk matches or slightly exceeds Run #17 at all packet sizes and rates. The GUE Option::is_some() check in send/recv hot paths adds zero measurable overhead when GUE is not configured. https://claude.ai/code/session_01NmhvqXvaxmA3fSXRGrzPpb --------- Co-authored-by: Claude <noreply@anthropic.com>
gspivey
added a commit
that referenced
this pull request
Apr 17, 2026
* feat: GUE tunnel endpoint (Generic UDP Encapsulation) Implement RFC 8470-style L3-over-UDP encapsulation as a transparent tunnel mode on UdpSocket. When GUE is configured, the application's send_to/recv_from calls are automatically encapsulated/decapsulated through the tunnel. Wire format: [Outer Eth][Outer IPv4][Outer UDP:6080][GUE 4B][Inner IPv4][Inner UDP][Payload] Key changes: - New dpdk-udp/src/gue.rs: GUE header build/parse, frame encap/decap functions - UdpSocket.set_gue()/gue(): per-socket tunnel configuration - TX path: builds GUE-encapsulated frames, ARP resolves tunnel remote endpoint - RX path: decapsulates matching GUE frames, returns inner src address to app - NetworkConfig.with_gue(): builder integration parallel to VLAN - MTU check accounts for 32-byte GUE overhead (outer IP + outer UDP + GUE header) - 23 unit tests: header codec, frame roundtrip, socket-level decap, PPS benchmark - README roadmap updated: GUE moved from Planned to Done https://claude.ai/code/session_01NmhvqXvaxmA3fSXRGrzPpb * docs: add Run #19 perf results — GUE tunnel endpoint regression check No regressions detected. rust-dpdk matches or slightly exceeds Run #17 at all packet sizes and rates. The GUE Option::is_some() check in send/recv hot paths adds zero measurable overhead when GUE is not configured. https://claude.ai/code/session_01NmhvqXvaxmA3fSXRGrzPpb --------- Co-authored-by: Claude <noreply@anthropic.com>
gspivey
pushed a commit
that referenced
this pull request
May 20, 2026
gspivey
pushed a commit
that referenced
this pull request
May 20, 2026
Implement ICMPv6 error message parsing and handling (IPv6 roadmap task 8): - Destination Unreachable (type 1): no route, admin prohibited, beyond scope, address unreachable, port unreachable - Packet Too Big (type 2): carries Next-Hop MTU - Time Exceeded (type 3): hop limit exceeded, fragment reassembly - Parameter Problem (type 4): erroneous header, unrecognized next header Adds Icmpv6ErrorInfo, parse_icmpv6_error(), Icmpv6Action enum, process_icmpv6_full() on Icmpv6Handler, and RX path integration via the existing error_queue/take_error() infrastructure. 24 new tests, no performance regression (Run #17 in perf-test-log.md).
gspivey
added a commit
that referenced
this pull request
Jun 13, 2026
## Roadmap Item Implements **roadmap item #17**: `dpdk-stdlib-tcp`: TimerWheel, CongestionState, and Tcb. Spec: `.kiro/specs/tcp-support/` tasks 5.5, 5.6, 5.7. ## Changes ### `dpdk-stdlib-tcp/src/timer.rs` — TimerWheel - 1ms granularity timer wheel for TCP timer management - 6 timer types: RTO, Persist, Keepalive, TimeWait, FinWait2, DelayedAck - Operations: insert, cancel, cancel_all, expired (tick-advance), next_deadline, is_active - Each connection can have at most one active timer per type - 8 unit tests covering insert/expire, cancel, replace, multiple timers, all types ### `dpdk-stdlib-tcp/src/congestion.rs` — CongestionState - RFC 5681 congestion control: slow-start + congestion avoidance - RFC 6298 RTT/RTO estimation: α=1/8, β=1/4, Karn's algorithm, clamped [1s, 60s] - Initial window per RFC 6928: min(10×MSS, max(2×MSS, 14600)) - Fast retransmit/recovery (NewReno): on_triple_dup_ack, on_partial_ack, on_recovery_exit - effective_window(rwnd), backoff_rto, on_rto - 19 unit tests covering all operations and edge cases ### `dpdk-stdlib-tcp/src/tcb.rs` — Tcb struct - Full Transmission Control Block with all per-connection state - Send/receive sequence variables (RFC 9293 §3.3.1) - Window scaling, MSS (local + peer) - CongestionState, timer deadlines, retransmit_count - Engine-internal buffers: send_buf, retransmit_queue, reorder_buffer - Nagle state, socket options, MAC addresses, shared ConnectionHandle - Helper methods: effective_mss, flight_size, available_send_window - 5 unit tests ## Tests Added - 8 timer wheel tests - 19 congestion control tests - 5 TCB tests - All 88 unit tests + 10 property tests pass for dpdk-stdlib-tcp - Full workspace: 0 failures ## Tradeoffs - TimerWheel uses HashMap internally (O(n) expiration scan) rather than a hierarchical wheel with slots. Sufficient for MVP connection counts; can be optimized later if needed for thousands of concurrent connections. - CongestionState::on_ack takes `bytes_acked` parameter (prefixed unused for now) — the slow-start/CA algorithms use per-ACK increment rather than bytes-based, but the parameter is reserved for future use in RFC 3465 (ABC) or CUBIC. --------- Co-authored-by: Agent Router <agent@agent-router.dev>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Add a complete performance benchmarking pipeline that measures UDP echo
throughput, latency, and packet loss across 4 configurations:
Infrastructure:
Orchestration (scripts/run-perf-tests.sh):
New app: apps/plain-echo — minimal std::net UDP echo (~30 lines, no deps
beyond clap) for the kernel baseline measurement.
https://claude.ai/code/session_01KQw8H7AZWpb5jQLSNpPiGr