Repository navigation
feat: jumbo frames, checksum offload, and bindgen compatibility - #32
Merged
Merged
Conversation
- Add 8500 to default TREX packet sizes (64,512,1400,8500) - Increase echo app recv buffers from 2048 to 10000 bytes for jumbo payloads - Set MTU 9001 on DUT kernel interfaces for jumbo frame support - Add --max-pkt-len=9100 --mbuf-size=10240 to testpmd for jumbo frames - Add --mbuf-factor 8 to TRex startup for large packet memory allocation https://claude.ai/code/session_01294ECCWxYiENcnK5HVVyoL
…ling - Add mbuf_9k: 4096 to trex_cfg.yaml memory config so TRex can allocate mbufs large enough for 8500B packets - Remove --mbuf-factor 8 (explicit pool config is the proper fix) - Add try/except per packet size in run_benchmark.py so one size failing doesn't crash the entire benchmark — partial results are still written https://claude.ai/code/session_01294ECCWxYiENcnK5HVVyoL
- Set DPDK port MTU to 9001 (AWS VPC max) in both the UdpSocket and DpdkBackend paths so the NIC accepts jumbo frames - Increase mempool data_room_size from 2176 to 9344 bytes (9216 + headroom) so mbufs can hold full jumbo frames without truncation - Use force=True in TRex client.start() to bypass the 16 Gbps line rate check — ENA PMD falsely reports 16 Gbps but c5n actually has 25 Gbps Root cause of 100% drop for rust-dpdk 8500B: the DPDK port's default MTU was 1500, so the NIC silently dropped all jumbo frames on RX. https://claude.ai/code/session_01294ECCWxYiENcnK5HVVyoL
- Change instance type from c5n.2xlarge to c6in.xlarge - Remove force=True from TRex to see if c6in ENA reports accurate line rate - c6in.xlarge: 4 vCPUs, 8 GB, 30 Gbps — $0.284/hr vs $0.432/hr https://claude.ai/code/session_01294ECCWxYiENcnK5HVVyoL
- Move try/except to per-rate-step so partial results (e.g. 70K, 140K succeed but 350K hits line rate) are still saved to the JSON - Sort packet sizes numerically in the summary table (64B < 512B < 1400B < 8500B) instead of lexicographic string sort https://claude.ai/code/session_01294ECCWxYiENcnK5HVVyoL
Stubs define RTE_PKTMBUF_HEADROOM as u16 but bindgen generates a wider type. The 9216 + HEADROOM expression was u32 on EC2, causing a type mismatch with with_data_room_size(u16). https://claude.ai/code/session_01294ECCWxYiENcnK5HVVyoL
Define JUMBO_DATA_ROOM_SIZE as a u16 constant (9344 = 9216 + 128) with a const assertion that catches overflow at compile time. Removes the `as u16` cast that could silently truncate. https://claude.ai/code/session_01294ECCWxYiENcnK5HVVyoL
… IMDS The deploy message, results summary, and report JSON all had hardcoded "c5n.2xlarge". Now queries the actual instance type from IMDS after SSM is ready, so it correctly reflects whatever instance type the CDK stack deploys. https://claude.ai/code/session_01294ECCWxYiENcnK5HVVyoL
Query the port's reported line rate from TRex and skip rate steps that would exceed 95% of it. This avoids the "Expected L1 B/W exceeds port line rate" error without using force=True. For 8500B packets on ENA (reports 16 Gbps): max ~228K PPS. Rate steps above that are skipped with a clear message. https://claude.ai/code/session_01294ECCWxYiENcnK5HVVyoL
TRex benchmark: - Use force=True to bypass ENA's false 16 Gbps line rate report - Cap PPS for jumbo frames (>1500B) to stay under 30 Gbps instance bandwidth — standard packets keep all rate steps uncapped - Configurable via --bandwidth-cap-gbps (default 30) DPDK jumbo echo fix: - Auto-detect interface MTU from /sys/class/net/<iface>/mtu during routing table initialization. On AWS ENA this reads 9001, so send_to() no longer rejects jumbo payloads as "too large" - Also parse MTU column from /proc/net/route as a first source Root cause of rust-dpdk 100% drop at 8500B: the routing table defaulted to MTU 1500, so send_to() rejected the 8458-byte echo payload. https://claude.ai/code/session_01294ECCWxYiENcnK5HVVyoL
Integration tests: - Add --payload-size flag to test-client for binary jumbo payloads - Increase test-client recv buffer from 1024 to 10000 bytes - Add tier1 test 5 (jumbo_diagnostics): dumps interface MTU, route table MTU column, and echo server config to CI logs - Add tier1 test 6 (jumbo_echo_8000): sends 8000B payload, verifies echoed response matches size Echo server: - Print MTU and max_udp_payload at startup for diagnostics Routing: - Auto-detect interface MTU from /sys/class/net/<iface>/mtu - Parse MTU column from /proc/net/route Perf tests: - Fix instance_type in results (use env var, not unexpanded shell var) - Cap jumbo PPS to 30 Gbps bandwidth limit, use force=True for TRex https://claude.ai/code/session_01294ECCWxYiENcnK5HVVyoL
Root cause: when the DPDK ENI is bound to vfio-pci, it has no kernel interface, so auto_detect_routing() can't read its MTU from sysfs or /proc/net/route. The routing table defaults to MTU 1500, causing send_to() to reject jumbo payloads with "Payload too large: max 1472". Fix: after auto-detection, if we're on a DPDK backend, override the routing table MTU to 9001 to match the port's rxmode.mtu configuration. Confirmed by integration test diagnostics: - Echo server (receiver): MTU=9001, max_udp_payload=8973 ← correct - Test client (sender): "Payload too large: max 1472, got 8000" ← bug https://claude.ai/code/session_01294ECCWxYiENcnK5HVVyoL
The three build_udp_* functions had a hardcoded MAX_UDP_PAYLOAD (1472) check that rejected jumbo payloads even when the routing table allowed them. Changed to use MAX_FRAME_SIZE - TOTAL_HEADER_LEN (8973 bytes) as the absolute maximum, matching the jumbo MTU 9001 configuration. The routing table's max_udp_payload() in send_to_addr() remains the authoritative per-socket MTU guard; the build functions are now just a safety net for the absolute frame size limit. https://claude.ai/code/session_01294ECCWxYiENcnK5HVVyoL
Enable ENA Express (Scalable Reliable Datagram) on all data plane ENIs in the perf test CDK stack. This raises the single-flow UDP throughput limit from 5 Gbps to 25 Gbps on c6in instances. Previous results showed rust-dpdk and native-dpdk hitting the exact same ~5.3 Gbps ceiling at 8500B — confirmed as ENA's single-flow cap. Also add a placeholder workflow for future multi-flow perf tests. https://claude.ai/code/session_01294ECCWxYiENcnK5HVVyoL
… it) The CfnNetworkInterfaceProps type doesn't have enaSrdSpecification yet. Use addPropertyOverride to set EnaSrdSpecification directly on the CloudFormation resource. Verified via cdk synth — all 3 data plane ENIs (TRex TX, TRex RX, DUT) have ENA Express enabled. https://claude.ai/code/session_01294ECCWxYiENcnK5HVVyoL
…tion CloudFormation's AWS::EC2::NetworkInterface doesn't support EnaSrdSpecification. Instead, enable ENA Express after stack deploy using `aws ec2 modify-network-interface-attribute --ena-srd-specification` on all 3 data plane ENIs (TRex TX, TRex RX, DUT). https://claude.ai/code/session_01294ECCWxYiENcnK5HVVyoL
c6in.xlarge had only 6.25 Gbps baseline — sustained 30s tests exhausted burst credits and settled to baseline. c6in.8xlarge has 50 Gbps sustained baseline with ENA Express support, enough to show the real jumbo frame throughput advantage. Also bump bandwidth cap from 30 to 50 Gbps to match. https://claude.ai/code/session_01294ECCWxYiENcnK5HVVyoL
ENA Express (SRD) requires MTU <= 8900 bytes per AWS documentation (amzn-ec2-ena-utilities/check-ena-express-settings.sh). Our jumbo MTU 9001 exceeds this limit, which caused 90%+ packet drops on c6in.8xlarge. Reverting to c6in.xlarge without ENA Express: - Remove ENA Express post-deploy enablement from run-perf-tests.sh - Revert instance type from c6in.8xlarge to c6in.xlarge - Revert bandwidth cap from 50 to 30 Gbps Proven single-flow results on c6in.xlarge: - rust-dpdk matches native-dpdk within 1% at all packet sizes - 8500B jumbo: ~5.3 Gbps (ENA single-flow cap, not our stack) - 1400B: ~7.8 Gbps at 700K PPS with <0.5% drop Future work for higher bandwidth: - Multi-flow tests (separate workflow placeholder created) - ENA Express with MTU <= 8900 on c6in.8xlarge+ https://claude.ai/code/session_01294ECCWxYiENcnK5HVVyoL
RX path: verify IPv4 header and UDP checksums in software on every received packet. Corrupted packets are dropped before reaching the application. Handles UDP checksum-disabled (0) per RFC 768. Validation runs in both the inline recv path and the pipeline worker loop. TX path: when the NIC supports hardware checksum offload (e.g., ENA), the DPDK backend sets mbuf ol_flags (RTE_MBUF_F_TX_IP_CKSUM, RTE_MBUF_F_TX_UDP_CKSUM) and writes a pseudo-header checksum so the NIC computes final checksums. Falls back to software checksums on NICs without offload or non-DPDK backends. Changes: - dpdk-sys: add RTE_MBUF_F_TX/RX offload flag constants - dpdk: add ol_flags/tx_offload accessors on Mbuf wrapper - dpdk-udp: add verify_ipv4_checksum, verify_udp_checksum, udp_pseudo_header_checksum functions; enable checksum offloads in port config; apply TX offload in DPDK send path; validate RX checksums in process_frame_zerocopy and worker_loop - 10 new unit tests covering validation, corruption, and offload fields https://claude.ai/code/session_01FDpfLChZCtgFqko3oX96e8
…-jumbo-frames-rZQcl # Conflicts: # dpdk-udp/src/lib.rs
In real DPDK, tx_offload lives inside an anonymous union in rte_mbuf. Bindgen represents this as __bindgen_anon_N.tx_offload rather than a direct field, so (*mbuf).tx_offload fails to compile with real DPDK. Add dpdk_shim_set/get_mbuf_tx_offload() C wrappers and corresponding Rust shim/stub functions so the access is portable across both paths. https://claude.ai/code/session_01FDpfLChZCtgFqko3oX96e8
These are #define macros in rte_mbuf_core.h that bindgen cannot capture. Without them, the TX offload and RX validation code fails to compile when building with --features dpdk-sys/bindgen on real DPDK systems. https://claude.ai/code/session_01FDpfLChZCtgFqko3oX96e8
The test was reading tx_offload via direct field access which fails with real DPDK bindgen bindings. Use mbuf_get_tx_offload() shim instead. https://claude.ai/code/session_01FDpfLChZCtgFqko3oX96e8
[CI] Stage: DeployInfrastructure ready.
|
[CI] Stage: SummaryAll tests PASSED. ARP seeding: kernel /proc/net/arp (automatic)
|
✅ Integration Tests Passed (Run 24271256698)Branch: Test Results
Application Logs (last 20 lines)receiver-echo-server.log sender-echo-server.log sender-test-client.log receiver-test-client-iperf.log sender-test-client-iperf.log Full Application Logs (last 200 lines each)receiver-echo-server.logsender-echo-server.logsender-test-client.logreceiver-test-client-iperf.logsender-test-client-iperf.log
|
2 of 4 tasks
gspivey
pushed a commit
that referenced
this pull request
Apr 12, 2026
…test Add full 802.1Q VLAN support to the socket layer: - VlanConfig type with VLAN ID (0-4094), PCP priority (0-7), DEI - build_udp_frame_into_vlan() for VLAN-tagged frame construction - parse_udp_packet/parse_udp_packet_ref transparently handle VLAN tags - verify_ipv4_checksum/verify_udp_checksum work on VLAN-tagged frames - process_frame_zerocopy dispatches ARP/ICMP/UDP correctly for tagged frames - ARP and ICMP parsers handle VLAN-tagged frames via detect_vlan() helper - Per-socket VLAN config via set_vlan()/vlan() and NetworkConfig::with_vlan() - UdpSocketBuilder applies VLAN config from NetworkConfig - 16 new unit tests covering build/parse/checksum/roundtrip/config Also adds Tier 4 jumbo frame echo integration test that exercises the jumbo frame feature (PR #32) with 1400/4000/8000-byte payloads on EC2. https://claude.ai/code/session_01Tumf1bXMixEaMKzgLvcBbD
2 of 4 tasks
gspivey
pushed a commit
that referenced
this pull request
Jun 18, 2026
gspivey
added a commit
that referenced
this pull request
Jun 20, 2026
## ROADMAP Item #32 Implements the TRex TCP performance profile and benchmark runner (tasks 14.1–14.4 from `.kiro/specs/tcp-support/tasks.md`). ### Changes 1. **`scripts/perf-tests/trex/tcp_echo_profile.py`** — TRex ASTF (Advanced Stateful) TCP echo profile. Generates TCP connections: client sends payload, server echoes, teardown. Configurable payload size via tunables. 2. **`scripts/perf-tests/trex/run_tcp_benchmark.py`** — TCP benchmark runner covering four payload sizes (64/512/1400/65536 B). Collects: - P50/P90/P99 latency percentiles - CPS (connections per second) - Throughput (Mbps) - Retransmits, timeouts, connection drops - Structured JSON output: `test_name, backend, metric_name, metric_value, unit` 3. **`apps/plain-tcp-echo`** — Minimal kernel TCP echo server using `std::net::TcpListener/TcpStream`. DUT baseline config (`plain-rust-tcp`) for comparison against DPDK TCP paths. 4. **`scripts/perf-tests/tcp-dut-configs.md`** — Documents all TCP DUT configurations (plain-rust-tcp, rust-dpdk-tcp, tokio-dpdk-tcp) and benchmark parameters. ### Testing - `cargo check --workspace` — passes - `cargo test -p plain-tcp-echo` — passes (binary crate, no unit tests) - `cargo test -p dpdk-stdlib-tcp` — all existing TCP tests pass (no regressions) ### Tradeoffs - TCP requires TRex ASTF mode (stateful) unlike UDP which uses STL (stateless). The profile and runner are separate from the UDP infrastructure because TRex exposes different APIs/metrics for ASTF. - The `plain-tcp-echo` uses thread-per-connection with 64KB read buffer — matching typical kernel server patterns for a fair baseline comparison. --------- Co-authored-by: Agent Router <agent@agent-router.dev>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
ol_flagsandtx_offloadfields are set with proper pseudo-header checksums. Software fallback remains for NICs without offload support.verify_ipv4_checksum()andverify_udp_checksum()functions for software verification of received packets.udp_pseudo_header_checksum()helper for TX offload path.tx_offloadfield lives inside an anonymous union inrte_mbuf, so direct field access breaks with real DPDK bindgen bindings. Added C shim functions (dpdk_shim_set_mbuf_tx_offload/dpdk_shim_get_mbuf_tx_offload) and corresponding Rust wrappers, plus allRTE_MBUF_F_TX/RXconstants that bindgen can't capture from#definemacros.Files changed
dpdk-sys/csrc/dpdk_shim.c— C shim fortx_offloadfield accessdpdk-sys/src/shim.rs— Rust wrappers + mbuf offload constants (real DPDK path)dpdk-sys/src/stubs.rs— Matching stubs (no-DPDK path)dpdk/src/mbuf.rs—set_tx_offload()uses shim, new offload accessorsdpdk-udp/src/lib.rs— Checksum functions, TX offload in send path, jumbo payload limitsdpdk-udp/src/routing.rs— MTU override for DPDK backendsapps/test-client/src/main.rs— Jumbo frame CLI supportscripts/— Integration and perf test updatesTest plan
cargo build && cargo test)https://claude.ai/code/session_01FDpfLChZCtgFqko3oX96e8