Skip to main content
KVIZK.1
Associate II
October 4, 2021
Question

Byte is lost from UART packet after being put into memory (DMA)

  • October 4, 2021
  • 30 replies
  • 7960 views

We have run into a problem where a new run of boards with STM32F429ZIT MCU's have been losing a byte (always loses a single byte which seems odd) from the UART communications that are sent to DMA.

We send packets of UART data to the MCU and the data is then sent to DMA. When we read the data back from DMA we will occasionally read back a packet with a missing byte. The failure rate is about 1 in 10 Megabytes or 1 in 50,000 packets.

I have scoped the UART communications during the point of failure and the failed data packet looks perfect at the pin (so UART signal is ruled out) and there are no transients on VDD or 3V3. Several people have looked at the design and have not determined anything obviously wrong (though of course something is)

We have solved this issue by using retry logic in software, but management is demanding an answer as to why the new board run is having this issue.

So I have two questions:

1: Has anyone encountered an issue like this and have any advice on what could be the root cause?

2: From your experience, is such a failure rate of 1 in 10 Megabytes or 1 in 50,000 packets considered a normal or expected problem with these kinds of systems? In other words, did we just get lucky in the past with not needing retry logic for this kind of system?

Thanks for the help!

This topic has been closed for replies.

30 replies

TDK
October 4, 2021

You shouldn't see failures like this. Are you sure there's not a race condition in the code? How are you receiving data? DMA with always-receiving circular buffer or is the UART inactive between receptions? Is the missing byte always the first byte?

"If you feel a post has answered your question, please click ""Accept as Solution""."
KVIZK.1
KVIZK.1Author
Associate II
October 4, 2021

UART data from the RX pin (pin 102) is run directly into DMA process, my understanding is that because of this we have no visibility until we read the data from memory. There certainly could be a race condition, but the problem is that the software is the same between board runs and the previous board does not fail. The byte lost is at a random position in the packet. Packet size is typically 172 bytes and we lose it somewhere in the middle.

waclawek.jan
Super User
October 4, 2021

What do you mean by "missing byte" exactly? Post example.

How exactly are you using the DMA? Post code.

Which USART, what baudrate, which DMA, what is the system clock, what other peripherals use DMA?

Do you have an USART errors handler in place? Are there USART overruns?

JW

KVIZK.1
KVIZK.1Author
Associate II
October 4, 2021

"What do you mean by "missing byte" exactly? Post example."

We have a packet that looks like this in the logic analyzer:

0693W00000FBOWIQA5.pngThe first byte 0xAB is the packet size of 171 (actual packet size is 172 with size byte included). When this packet is read from the DMA we will get all of the same data less 1 byte somewhere in the middle.

"How exactly are you using the DMA? Post code."

 Dma2Init(); is called once in main.c. I believe this is interrupt driven. Note that we have not had an issue with this code on past boards.

#define DMA2_C
 
#include <stdint.h>
#include <stdbool.h>
#include <string.h>
 
 
// CMSIS includes
//
#include <stm32f429xx.h>
 
 
// application includes
//
#include "qhal_uart.h"
#include "Dma2.h"
 
 
// Module function prototypes
//
void DMA2_Stream2_IRQHandler(void);
void Dma2Init(void);
 
 
// Local module variables
//
uint32_t dma2Stream2Error;
 
 
// Privately imported functions and variables
//
extern uint8_t us1_rxbuf[US1BUFS_LEN];
 
 
 
void DMA2_Stream2_IRQHandler(void)
/*
// Handles enabled interrupts from DMA2, Stream2
//
//
// Receives: not applicable to an ISR
//
// returns: not applicable to an ISR
*/
{
uint32_t dma2Irq;
 
 dma2Irq = (DMA2->LISR & DMA2_S2_IRQ_MASK);
 
 // if one of the 2 buffers has been filled reset its target address
 if((dma2Irq & DMA_LISR_TCIF2) > 0)
 {
 // if the current target buffer is 1 reset target zero's address
 if((DMA2_Stream2->CR & DMA_SxCR_CT) > 0)
 DMA2_Stream2->M0AR = ((uint32_t) &us1_rxbuf[0]);
 
 // else the current target buffer is 0 so reset target one's address
 else
 DMA2_Stream2->M1AR = ((uint32_t) &us1_rxbuf[HALF_MAX_US1BUFS]);
 }
 
 // if DMA errors occurred flag them for the application to handle
 if((dma2Irq & (DMA_LISR_TEIF2 | DMA_LISR_DMEIF2)) > 0)
 dma2Stream2Error = dma2Irq;
 
 // clear the interrupts
 DMA2->LIFCR = dma2Irq;
}
 
 
 
void Dma2Init(void)
/*
// Initializes DMA2 to stream USART1 Rx data directly from USART1 to the Rx
// buffer.
//
// Configures Stream 2 to use channel 4 (USART1 RX), single transfers of
// 8 bits in size from USART1 to memory in double buffer mode
//
// Receives: void
//
// returns: void
*/
{
 // The initial write of DMA2_Stream2->CR to zero sets the following:
 //
 // Stream Channel 0
 // MBURST: Single transfer
 // PBURST: Single transfer
 // CT 0, current target is buffer 0
 // PL 0, lowest stream priority level
 // DBM: Single buffer mode
 // MSIZE: 8 bits
 // PSIZE: 8 bits
 // MINC 0, memory pointer is fixed
 // PINC 0, peripheral pointer is fixed
 // CIRC 0, circular mode disabled
 // DIR 0, Peripheral to Memory
 // PFCTRL 0, DMA controller as flow controller
 // xxIE 0, No interrupts enabled
 // EN 0, Stream is disabled
 //
 // The required changes are made after stream disable is confirmed
 DMA2_Stream2->CR = 0;
 
 while(DMA2_Stream2->CR & 1) // wait here for EN bit to go low
 ;
 
 DMA2->LISR = 0; // reset transfer status bits
 DMA2->HISR = 0;
 
 // set the peripheral port register address (USART1_DR)
 DMA2_Stream2->PAR = (USART1_BASE + 4);
 
 // set the memory buffer address
 DMA2_Stream2->M0AR = (uint32_t) (&us1_rxbuf[0]);
 DMA2_Stream2->M1AR = (uint32_t) (&us1_rxbuf[HALF_MAX_US1BUFS]);
 
 // number of data to be transfered
 DMA2_Stream2->NDTR = HALF_MAX_US1BUFS;
 
 DMA2_Stream2->CR = (4 << 25); // set stream2 to use channel 4
 DMA2_Stream2->CR |= (3 << 16); // set stream2 priority to very high
 
 // use direct mode not FIFO mode
 DMA2_Stream2->FCR = 0;
 
 // stream configuration
 DMA2_Stream2->CR |= DMA_SxCR_DBM; // enable stream2 double buffer mode
 DMA2_Stream2->CR |= DMA_SxCR_MINC; // memory increment mode
 DMA2_Stream2->CR |= DMA_SxCR_TCIE; // transfer complete interrupt enable
 DMA2_Stream2->CR |= DMA_SxCR_EN; // enable stream 2
 
 // reset DMA2, Stream2 error indicator
 dma2Stream2Error = 0;
 
 NVIC_SetPriority(DMA2_Stream2_IRQn, 10); // set DMA2 Stream2's IRQ priority
 NVIC_EnableIRQ(DMA2_Stream2_IRQn); // and unmask it
}
 

"Which USART, what baudrate, which DMA, what is the system clock, what other peripherals use DMA?"

 USART1 (Pin 120 for RX and pin 136 for TX on LQFP 144 pin package). DMA2 presumably. 2 Megabaud. System clock 8MHz, and I don't think we use DMA for anything else since there are no other references to it.

"Do you have an USART errors handler in place? Are there USART overruns?"

We did not have any error handlers in place until now where we implemented retry logic. When a packet read from DMA does not equal the size specified by the packet size byte we retry. It's possible there are overruns I have not checked.

Tesla DeLorean
Guru
October 4, 2021

>>Note that we have not had an issue with this code on past boards.

Don't seem to use volatile variables where appropriate.

Doesn't seem any useful reason to use double buffers to describe continuous memory buffers.

Should be using a circular mode.

No indication how large the buffer is, and the relationship between the size, and the failure points in the data stream.

>>2 Megabaud. System clock 8MHz

Unlikely to get 2Mbaud from 8 MHz bus clocks.

Tips, Buy me a coffee, or three.. PayPal Venmo (See Profile) Up vote any posts that you find helpful, it shows what's working..
waclawek.jan
Super User
October 4, 2021

> The first byte 0xAB is the packet size of 171 (actual packet size is 172 with size byte included). When this packet is read from the DMA we will get all of the same data less 1 byte somewhere in the middle.

How do you know where is the end of the packet? Is that byte missing, or just has an unexpected value?

> 2 Megabaud. System clock 8MHz,

8MHz, really? That would leave only 80 system cycles to process a Rx byte, I wouldn't be surprised if this would be the problem.

What are values of HALF_MAX_US1BUFS/HALF_MAX_US1BUFS?

How does "main" pick the values from the buffer?

> It's possible there are overruns I have not checked.

Then check them (I mean, set up USART interrupt and enable it only for the errors). It may be revealing.

JW

(Btw. not setting DMA to Circular mode is somewhat confusing, but the Double-Buffer mode enables Circular in hardware.)

TDK
October 4, 2021

If the missed byte is midstream, probably not a synchronization issue.

It could be a clock issue. 2 MBaud is workable, but the tolerance between chips needs to be adequate. Expect issues if you're using HSI.

I use the STM32F4 UART at ~5 MBaud and never had issues like this, even with significant amounts of data.

Look at the NF and FE bits in USARTx_SR which will be set if the clock is mismatched.

"If you feel a post has answered your question, please click ""Accept as Solution""."
KVIZK.1
KVIZK.1Author
Associate II
October 4, 2021

Can you elaborate on the NF and FE bits? I cannot find these in my source code. ty

TDK
October 4, 2021
See the reference manual for details.
*Bit 2 NF: Noise detected flag*
*This bit is set by hardware when noise is detected on a received frame.*
*Bit 1 FE: Framing error*
*This bit is set by hardware when a de-synchronization, excessive noise or
a break characteris detected.*
"If you feel a post has answered your question, please click ""Accept as Solution""."
waclawek.jan
Super User
October 4, 2021

0693W00000FBQPdQAP.png

KVIZK.1
KVIZK.1Author
Associate II
October 12, 2021

*Update

I mis-specified the clock speed. We have an external 8MHz oscillator that is for the real-time-clock and we use the internal clock to set CPU clock at 168MHz, sorry for the confusion.

I break-pointed at the point of communication failure and sure enough both of those bits are flipped:

Breakpoint set during normal operation

0693W00000FCGrPQAX.pngBreakpoint set at point of communication failure

0693W00000FCGuEQAX.pngI'm a little baffled by this since the UART data at the pin of the MCU looks perfectly fine when we get the failure. Is something going wrong inside of the MCU?

waclawek.jan
Super User
October 12, 2021

> Probably the clock is off a tiny amount. Perhaps your clock source is not as accurate as you think.

+1

Show us your clocks setup (relevant RCC registers content).

Also the clock frequency of the data source may be a little bit off.

JW

KVIZK.1
KVIZK.1Author
Associate II
October 12, 2021

Clock setup is as follows:

0693W00000FCHPgQAP.png 

There is a ton of data in RCC registers I don't know what is relevant.

0693W00000FCHX1QAP.png 

From the logic analyzer there is an occasional pulse width deviation on UART TX and RX of 0.8% (496-504ns for fastest bit). That's the most I've seen.

waclawek.jan
Super User
October 12, 2021

RCC_CFGR.SW (and SWS) to confirm that system clock runs off PLL; all fields of RCC_PLLCFGR to confirm PLL settings, mainly that it runs off HSE (and perhaps RCC_RCC_CR.HSEON/HSERDY to confirm HSE is up and running).

RCC_CFGR.PPRE2 to confirm that you've indeed set APB2 = AHB/4 = 42MHz. That's unusual (are you trying to spare power?). This setting also puts the baudrate divider to "strongly fractional" (not that it would be any better with 84MHz APB clock) 1.3125 (yes I should've noticed this on the USART registers' screenshot). With non-integer baudrate divisors, the requirements for precise baudrate matching become tighter. These things add up.

USART_SR.NE may (although not necessarily) also indicate that the edges are not as clean as they ought to be; you should perhaps look at them using an oscilloscope, and you should perhaps also review return/ground arrangements.

JW

KVIZK.1
KVIZK.1Author
Associate II
October 29, 2021

*Update

The source of the UART failure looks to be a transient on SYSCLK. I outputted SYSCLK onto MCO2 and put a scope on the pin and have been monitoring the signal. This is SYSCLK (168MHz) downscaled by 4x to 42MHz. At the point of the communication failure this is what we see

0693W00000GVvizQAD.png 

During this transient the frequency of the clock will dip down from 42MHz to as low as 20 MHz with quite a bit of noise. These transients occur anywhere from several times a minute to a few times a day.

Note that there is no device connected to MCO2 during this test, MCO2 pin is open.

Here is SYSCLK normally as seen on pin MCO2

0693W00000GVvdqQAD.png 

As the glitch occurs the SYSCLK signal begins to breakdown and form this shape. It appears that there is another sinewave at about 1/3 frequency riding on the SYSCLK signal. The amplitude and frequency of SYSCLK are also changed

0693W00000GVvk2QAD.png 

I’ve monitored the HSE clock on the crystal and outputted HSE on MCO1 and do not see any transients on HSE during the UART failure or this clock glitch. I also do not see any transients on the 3.3V bus during these failures. So again I am scratching my head on the cause!

TDK
October 29, 2021

Well, that would certainly explain the results you're seeing. Thanks for posting those. Your clock settings above look fine to me. I haven't seen anything like this. I can't think of an explanation that is consistent with HSE not being interrupted.

Please post if you solve it. Does it happen on all boards or just some?

"If you feel a post has answered your question, please click ""Accept as Solution""."
KVIZK.1
KVIZK.1Author
Associate II
November 2, 2021

Manipulating the PLL settings affects the rate of failure. I will post my test results soon.

It happens on all of our new run of boards but I know of at least one "old" board that does it. The old boards have the same MCU/Memory/Crystal and have a slightly different layout. For the most part the old boards do not show the issue though. The only thing that jumps out is that the external oscillator on the older boards is placed a little more optimally. However when I monitor HSE it looks almost identical between known good and bad board.

waclawek.jan
Super User
October 29, 2021

Can you set up a simple continuous PWM output of day 1MHz on any of the timers and measure its frequency change during the event?

JW

KVIZK.1
KVIZK.1Author
Associate II
November 16, 2021

I have not done this test but I've confirmed that switching from HSE to HSI fixes the issue. So I've identified the HSE as the root cause.