DOCA Skills Cut BlueField-3 RDMA Code by 73%
In a BlueField-3 comparison, two coding agents built Go applications that used NVIDIA DOCA to send RDMA traffic and verify its arrival on the hardware. The demo argues that NVIDIA’s DOCA skills changed the route, not the result: both agents succeeded, but the agent with curated guidance reused NVIDIA’s sample and wrote about 73% less code while making roughly half as many hardware commands.

The skills changed the route, not the required result
Both coding agents had the same task: build a Go program that uses NVIDIA DOCA to send RDMA traffic on a BlueField-3, then verify that the bytes arrived. The prompt explicitly required DOCA. Both agents produced a working DOCA application with traffic verified on the network card.
The difference was whether one agent had access to NVIDIA’s DOCA skills. Without them, it had to find the relevant module, headers, sample code, and build process by probing the hardware and trying alternatives. With them, it opened the doca-rdma skill and was directed to NVIDIA’s shipped RDMA sample and the build details needed to use it.
That distinction matters because the skills did not relax the task or allow a shortcut around DOCA. They supplied platform-specific information that the unskilled agent had to rediscover: which module to link, how to configure Go’s C interoperability, where the sample lived, and which DOCA version to target.
The demonstrations report roughly 37 hardware commands without the skills and roughly 20 with them. Across the runs shown, the command counts were 30, 35, and 46 without skills, compared with 18, 20, and 22 with them. The outcome was the same; the skilled agent reached it with fewer trips to the hardware.
A sample and build recipe replaced trial and error
The doca-rdma skill pointed the agent to the doca pkg-config package, which pulls in libdoca_rdma, and gave it the Go cgo directive #cgo pkg-config: doca. It also identified NVIDIA’s installed sample directory, /opt/mellanox/doca/samples/doca_rdma/, including the doca_rdma_send and doca_rdma_receive examples.
Instead of improvising an implementation from scattered discoveries, the agent used the shipped sample as its starting point. The skills also specified that the application should use DOCA RDMA rather than fall back to raw libibverbs, and that it should be built against the DOCA version already installed on the card. A separate doca-programming-guide skill supplied the DOCA progress-engine task and event model.
Without that information, the other agent searched headers and probed the system to identify what was available, then worked out the build recipe as it went. Each discovery required another command on the hardware. The comparison presents the skills as a way to make the supported path explicit before implementation begins—not as a change to what “working” means.
Reusing NVIDIA’s sample cut handwritten code by 73%
The code comparison is substantial. The agent without skills wrote 695 lines of its own DOCA code. The agent with skills added a 189-line wrapper around NVIDIA’s shipped sample, a reported reduction of about 73% in handwritten code. Both applications used libdoca_rdma, and both verified traffic on the network card.
The reduction came from reusing NVIDIA’s implementation rather than rebuilding the low-level pieces from scratch. The demo attributes that change to the doca-rdma and programming-guide skills: one directed the agent to the module, sample, and build setup; the other supplied the programming model. The resulting application still had to send real RDMA traffic and demonstrate that the payload arrived.
Tests without an explicit DOCA requirement show the larger difference
The demo also describes comparisons in which the prompt did not mention DOCA. These tests make a different point from the side-by-side RDMA task, where both agents were already instructed to use the platform.
For flow, the agent without skills steered packets using a Linux kernel command and linked no DOCA libraries. With the skills, it built a DOCA flow pipeline and read hardware counters. For RDMA, the unskilled agent wrote 660 lines using raw libibverbs; the skilled agent used an 87-line wrapper around NVIDIA’s sample.
Across Flow, RDMA, Comm Channel, and Ethernet, the skills moved the agent toward DOCA. For DMA and SHA, the agent already used DOCA, and the demo says the skills did no harm. The presenters describe these as observed results across tests, not as win rates.
Taken together, the comparisons point to a practical advantage: platform-specific guidance can steer an agent toward NVIDIA’s supported samples and APIs, reducing both handwritten code and hardware probing. In the explicit-DOCA RDMA task, it did not make the difference between success and failure. It changed how much rediscovery was required to reach the same verified result.