2 minute read

Symptom

On the NUC, I219 was bound and the link was up — lsdev showed i219 net link=up. A ping still died on “descriptor written, hardware never completes”: TDT advanced, TDH stayed put, last_rc=-3. Serial tx_ok stayed at 0. Class path stuck at “see the NIC, cannot send.”

First hypothesis

I started with “registers not enabled enough”: maybe I219 needs extra TXDCTL / TARC0 gates. On many Intel GbE variants those bits look like switches. Knife one went straight at “write the gates and it will move.”

Tries and failures

TXDCTL written — still -3. Add TARC0 — same. Gate hypothesis out, at least as the root cause.

Then a dummy TX flush after the ring was ready. Serial finally printed Flush OK — hardware can give Descriptor Done. Formal Send was still -3. Negated: “hardware is totally dead / never produces DD.” Not negated: “the formal path wrecks TX state somewhere.”

A string of falsifications followed: keep FEXTNVM11 after flush — still dead; move flush after link-up — Flush NoDD; full TCTL then flush — flush dies; disable MSI / poll — flush still broken (MSI not the main cause); “EN-only for flush, then full TCTL” — not enough; flush right after ring base is programmed — Flush OK again, formal Send still -3.

The picture was awkward: early flush can OK; something before formal TX mutes the hardware again.

Turning point

The turn was shifting from “what else to write” to “what must not be rewritten after a good flush.” Failures lined up: any path that slapped a “full / old” TCTL over the chip after Flush tended to go dead; keeping only TCTL.EN after a successful flush finally moved the counters.

Root cause

Short version: rewriting TCTL after the flush cleared a working TX state. The flush pushed the HW into a DD-capable state; the later “more correct looking” full TCTL write wiped it. Hence Flush OK vs formal Send forever -3.

After the fix, NUC tx_ok rose, static ping to the gateway replied, and clearing the netif for DHCP got an address. Class path connected.

Lessons

  • One hypothesis per knife; on failure write what you negated — do not spray the whole PHY/register family in one change.
  • Order matters: bind → MAC → TX counters/regs → minimal TX hypothesis; do not jump to real-machine DHCP before static ping.
  • “Write more registers” is not always safer; sometimes the right move is write less. Holding state after a good flush beats another “complete config” pass.
  • QEMU green is not real-HW progress; TX changes need NUC. Do not silently retry already-failed hypotheses.

Updated: