NetXDuo HTTPS downloads over 16 KB stall – root cause is a missing HAL_ETH_ErrorCallback in nx_stm32_eth_driver.c (fix included)
This is the follow-up to my 2023 thread, which is now closed for replies:
https://community.st.com/stm32cubeide-mcus-28/azurertos-netxduo-unable-to-download-16k-or-larger-image-using-web-http-client-secure-124990
Short version: the stall is in the ST NetX Ethernet glue driver (`nx_stm32_eth_driver.c`), not in NetX Duo, and not a packet-pool sizing problem. The fix is about 40 lines and still applies to the current `STMicroelectronics/stm32-mw-netxduo` main branch, which has the same driver code. I have filed it with ST as a bug with the full write-up and patch: https://github.com/STMicroelectronics/stm32-mw-netxduo/issues/1
**Symptom (from the original thread)**
HTTPS downloads of 16 KB or more via `nx_web_http_client` + NetX Secure stalled on every attempt on F417, F429 and H723. Plain HTTP downloaded multi-MB files fine with the same pool. After a stall the socket stayed ESTABLISHED and the pool showed fewer free packets than expected, which I wrongly reported as "free_packets corrupted". The only known workaround was a huge packet pool (128 × 1200 bytes or more).
**What actually goes wrong**
1. During a TLS download the HTTP client holds several packets while it reassembles a 16 KB TLS record. With a 4-descriptor RX ring and a modest pool, the pool briefly hits zero.
2. The HAL re-arms RX descriptors from `HAL_ETH_ReadData()` via `HAL_ETH_RxAllocateCallback()`, which calls `nx_packet_allocate(NX_NO_WAIT)`. While the pool is empty that allocation fails and the descriptor is left un-armed (`RxDescList.RxBuildDescCnt` stays non-zero).
3. Once every descriptor is un-armed, the DMA raises Receive Buffer Unavailable (RBU) and suspends. No further RX-complete interrupt can ever fire, so `HAL_ETH_ReadData()` is never called again and the ring is never re-armed, even after the pool frees up seconds later.
4. Result: a permanently dead receive path with an ESTABLISHED socket and a healthy pool. The "missing" packets were parked in the TLS reassembly chain and in the un-armed ring, waiting for data that could no longer arrive.
The HAL reports the RBU through `HAL_ETH_ErrorCallback()`, but the NetX glue driver never implements it. ST's own LwIP `ethernetif.c` does handle RBU; the NetX driver was simply missed. Plain HTTP never triggers it because the pool never goes to zero. That is why only HTTPS over about 16 KB failed, and why a very large pool "fixes" it.
**The fix (nx_stm32_eth_driver.c / .h)**
Handle RBU by re-running the receive path, and keep retrying from a 1-tick timer until the ring is fully re-armed. All deferred-event posting goes through one interrupt-safe helper because the timer runs outside the ETH ISR.
nx_stm32_eth_driver.h:
```c
#define NX_DRIVER_DEFERRED_PACKET_RECEIVED 1
#define NX_DRIVER_DEFERRED_DEVICE_RESET 2
#define NX_DRIVER_DEFERRED_PACKET_TRANSMITTED 4
#define NX_DRIVER_DEFERRED_RX_BUFFER_UNAVAILABLE 8
```
nx_stm32_eth_driver.c:
```c
static TX_TIMER nx_driver_rx_rearm_timer;
/* Called from the ETH ISR and from the ThreadX timer, so the RMW must be atomic. */
static VOID _nx_driver_post_deferred_event(ULONG event)
{
TX_INTERRUPT_SAVE_AREA
ULONG deferred_events;
TX_DISABLE
deferred_events = nx_driver_information.nx_driver_information_deferred_events;
nx_driver_information.nx_driver_information_deferred_events |= event;
TX_RESTORE
if (!deferred_events)
{
_nx_ip_driver_deferred_processing(nx_driver_information.nx_driver_information_ip_ptr);
}
}
static VOID _nx_driver_rx_rearm_timer_entry(ULONG id)
{
NX_PARAMETER_NOT_USED(id);
_nx_driver_post_deferred_event(NX_DRIVER_DEFERRED_RX_BUFFER_UNAVAILABLE);
}
/* In _nx_driver_initialize(), before the state is set to INITIALIZED: */
tx_timer_create(&nx_driver_rx_rearm_timer, "NetX ETH RX re-arm",
_nx_driver_rx_rearm_timer_entry, 0, 1, 0, TX_NO_ACTIVATE);
/* In _nx_driver_deferred_processing(), replacing the PACKET_RECEIVED block: */
if (deferred_events & (NX_DRIVER_DEFERRED_PACKET_RECEIVED | NX_DRIVER_DEFERRED_RX_BUFFER_UNAVAILABLE))
{
/* HAL_ETH_ReadData() also re-arms descriptors left un-armed by a failed allocation. */
_nx_driver_hardware_packet_received();
/* Pool was still empty when the HAL tried to re-arm: nothing else will retry, so we do.
tx_timer_change() is required to re-activate an expired one-shot timer. */
if (eth_handle.RxDescList.RxBuildDescCnt != 0U)
{
tx_timer_deactivate(&nx_driver_rx_rearm_timer);
tx_timer_change(&nx_driver_rx_rearm_timer, 1, 0);
tx_timer_activate(&nx_driver_rx_rearm_timer);
}
}
/* HAL callbacks: */
void HAL_ETH_RxCpltCallback(ETH_HandleTypeDef *heth)
{
NX_PARAMETER_NOT_USED(heth);
_nx_driver_post_deferred_event(NX_DRIVER_DEFERRED_PACKET_RECEIVED);
}
void HAL_ETH_TxCpltCallback(ETH_HandleTypeDef *heth)
{
NX_PARAMETER_NOT_USED(heth);
_nx_driver_post_deferred_event(NX_DRIVER_DEFERRED_PACKET_TRANSMITTED);
}
void HAL_ETH_ErrorCallback(ETH_HandleTypeDef *heth)
{
if ((HAL_ETH_GetDMAError(heth) & ETH_DMASR_RBUS) != 0U)
{
_nx_driver_post_deferred_event(NX_DRIVER_DEFERRED_RX_BUFFER_UNAVAILABLE);
}
}
```
Make sure `HAL_ETH_IRQHandler()` reaches the error callback: the RBU interrupt must be enabled in the DMA interrupt mask (it is in the default `ETH_DMAIER` setup of the F4 HAL) and `ETH_IRQn` must be enabled.
**Result**
STM32F417, STM32Cube FW_F4 1.28.3 (HAL 1.8.5), NetX Duo 6.5.2, 32-packet pool of 1536 bytes, HTTPS client window 8 × 1460:
- Before: 16 KB+ HTTPS downloads stalled on every attempt, socket ESTABLISHED, pool never recovered.
- After: 12 of 12 consecutive 1 MB (1,015,808 byte) HTTPS downloads, byte-exact and CRC-verified, 42–56 s each (17–23 KB/s). Each run logged 16–33 RBU events, all recovered within a few ticks.
So the big-pool workaround is not needed; it only hides the RBU by making the pool never run dry.
