HAL SPI nested interrupt race condition with overrun in IT mode transfer
Issue details
We have observed a situation where a full duplex HAL SPI transfer in IT mode would hang in permanent BUSY state after many transfers in a system with relatively long ISRs at higher priorities than the SPI.
Debugging showed that the lockup had the following premise:
- Transfer is started with HAL_SPI_TransmitReceive_IT
- A couple of bytes are transmitted
- HAL_SPI_IRQHandler executes with RXNE=1
- RxISR function pointer is called
- The SPI interrupt is pre-empted by a higher priority IRQ which takes long enough to result in RX FIFO overrun on the SPI peripheral (OVR=1)
- RxISR function continues and reads DR register
- HAL_SPI_IRQHandler returns (early return)
- HAL_SPI_IRQHandler executes with TXE=1
- HAL_SPI_IRQHandler reads SR register, clearing the OVR bit as a side effect (Reference manual: "Clearing the OVR bit is done by a read access to the SPI_DR register followed by a read access to the SPI_SR register")
- TxISR function pointer is called
- HAL_SPI_IRQHandler returns (early return)
- HAL_SPI_IRQHandler executes with RXNE=1
- HAL_SPI_IRQHandler reads SR but it does not contain the OVR error anymore
- Since the overrun went undetected, the HAL waits forever on the last RX interrupt that will never come
Product details
Tested with: STM32F0xx HAL Drivers v1.7.5
Also affects: all hal_spi implementations for SPI peripherals that clear OVR by sequentially reading DR and SR, including: STM32F4xx, STM32F7xx, STM32G0xx, STM32L4xx
Proposed fix
Removing the early returns in HAL_SPI_IRQHandler (in particular the one after TxISR()) prevents this issue, because then SR is checked for errors after TxISR is called without re-reading SR in between. Untested alternative: change the order from RX check -> TX check -> error check to RX check -> error check -> TX check.
