Skip to main content
Associate II
September 23, 2026
Solved

STM32H5E5ZJ Startup assembly Code raising core to VOS0 and System Clock to 248MHz

  • September 23, 2026
  • 6 replies
  • 99 views

I have written compact assembler code, with assistance from ChatGPT and Gemini, for the STM32H5E5ZJ to raise its core voltage to VOS0, configure PLL1 and PLL2 and switch system clock to 248MHz.

Any comments much appreciated as to the sequencing and safety of this code or any problems seen please.

It is tested and working reliably and occupies less than 512 Bytes. At the moment it runs in RAM and only after reset (not to be called later. To run in Flash it must also raise Flash wait states.  

The code also configures MCO1 and MCO2, configures SPI1 as master, running at 4MBaud, it puts SPI1 into a delayed loop sending test data readings from TIM2.

SPI1 is configured to use PA4 SPI1_NSS, PA6 SPI1_MISO, PA7 SPI1_MOSI and PG11 SPI1_SCK.

Any comments much appreciated as to the sequencing and safety of this code.

I have used ST Core MX extensively to help build the code by reading MX generated C code. Here is the code, anyone is free to make use of it as a supplement to official ST code if they so wish:-

 

.syntax unified; .cpu cortex-m33; .thumb

.equ PWR_BASE, 0x44020800; .equ VOSCR, 0x0010; .equ VOSSR, 0x0014
.equ RCC_BASE, 0x44020C00; .equ RCC_CR, 0x00; .equ RCC_CFGR1, 0x1C; .equ RCC_CFGR2, 0x20
.equ RCC_PLL1CFGR, 0x28; .equ RCC_PLL2CFGR, 0x2C; .equ RCC_PLL1DIVR, 0x34; .equ RCC_PLL2DIVR, 0x3C
.equ AFRL, 0x20; .equ AFRH, 0x24; .equ APB2ENR, 0xA4; .equ APB1LENR, 0x9C; .equ AHB2ENR, 0x8C; .equ CCIPR3, 0x0E0
.equ GPIO_BASE, 0x42020000; .equ GPIOA,0x0000; .equ GPIOB,0x0400; .equ GPIOC,0x0800; .equ GPIOD,0x0C00
.equ GPIOE,0x1000; .equ GPIOF,0x1400; .equ GPIOG,0x1800; .equ GPIOH,0x1C00; .equ MODER,0x00; .equ ODR,0x14
.equ SPI1_BASE, 0x40013000; .equ SPI1_CR1, 0x00; .equ SPI1_CR2, 0x04; .equ SPI1_CFG1, 0x08; .equ SPI1_CFG2, 0x0C
.equ SPI1_SR, 0x14; .equ SPI1_TXDR, 0x20
.equ TIM2_BASE, 0x40000000; .equ TIM2_CR1, 0x00; .equ TIM2_CNT, 0x24

.section .text; .word 0x20004000; .word init + 1; .global init; init:

@increase core voltage from vos3 0x00 to vos0 0x30
ldr r0,=PWR_BASE; ldr r1,=0x30; str r1,[r0,#VOSCR]
wait_vos: ldr r1,[r0,#VOSSR]; tst r1,#(1<<3); beq wait_vos


@enable GPIO A B C D E F G clocks
ldr r0,=RCC_BASE;
ldr r1,[r0,#AHB2ENR]; orr r1,r1,#0x7f; str r1,[r0,#AHB2ENR]
@enable TIM2 clock
ldr r1,[r0,#APB1LENR]; orr r1,r1,#(0x1<<0x00); str r1,[r0,#APB1LENR]
@enable SPI1 clock
ldr r1,[r0,#APB2ENR]; orr r1,r1,#(0x1<<12); str r1,[r0,#APB2ENR]
@select pll2_p as kernel clock for SPI1
ldr r1,[r0,#CCIPR3]; bic r1,r1,#(0x7<<0); orr r1,r1,#(0x1 << 0); str r1,[r0,#CCIPR3]
@route and scale PA8 MCO1 and PC9 MCO2 both /10 MCO1 is PLL1Q /10 MCO2 is PLL2P /10
ldr r1,[r0,#RCC_CFGR1]; ldr r3,=((0x07<<29)|(0x0f<<25)|(0x07<<22)|(0xf<<18)); bic r1,r1,r3
ldr r3,=((0x01<<29)|(0x0a<<25)|(0x03<<22)|(0x0a<<18)); orr r1,r1,r3; str r1,[r0,#RCC_CFGR1]
@set HSIDIV /1 PLL1 needs 64MHz clock
ldr r1, [r0,#RCC_CR]; bic r1,r1,#(0x3<<3); str r1,[r0,#RCC_CR]
@configure PLL1 for PLLM /4 N x31 P /2 Q /2 to give pll1p 248MHz
ldr r1,=0x0003040d; str r1,[r0,#RCC_PLL1CFGR]; ldr r1,=0x0101021e; str r1,[r0,#RCC_PLL1DIVR]
@configure PLL2 for PLLM /16 N x64 P /4 Q /4 R /4 to give pll2p 64MHz
ldr r1,=0x00071029; str r1,[r0,#RCC_PLL2CFGR]; ldr r1,=0x0303063f; str r1,[r0,#RCC_PLL2DIVR]
@ Enable PLL1 (RCC_CR Bit 24)
ldr r1,[r0,#RCC_CR]; orr r1,r1,#(1<<24); str r1,[r0,#RCC_CR]
wait_pll1: ldr r1,[r0,#RCC_CR]; tst r1,#(1<<25); beq wait_pll1
@ Enable PLL2 (RCC_CR Bit 26)
ldr r1,[r0,#RCC_CR]; orr r1,r1,#(1<<26); str r1,[r0,#RCC_CR]
wait_pll2: ldr r1,[r0,#RCC_CR]; tst r1,#(1<<27); beq wait_pll2
@switch sysclk to PLL1P
ldr r1,[r0,#RCC_CFGR1]; orr r1,r1,#(0x3<<0); str r1,[r0,#RCC_CFGR1]
wait_switch: ldr r1,[r0,#RCC_CFGR1]; and r1,r1,#(0x3<<3); cmp r1,#(0x3<<3); bne wait_switch

ldr r0,=GPIO_BASE; ldr r3,=GPIO_BASE+GPIOG
@GPIO pin configuration PA8 MCO1, PC9 MCO2 both AF0
ldr r1,[r0,#GPIOA+AFRH]; bic r1,r1,#(0x0f<<0); str r1,[r0,#GPIOA+AFRH]
ldr r1,[r0,#GPIOC+AFRH]; bic r1,r1,#(0x0f<<4); str r1,[r0,#GPIOC+AFRH]
@SPI1 pin configuration PA4 SPI1_NSS, PA6 SPI1_MISO, PA7 SPI1_MOSI, PG11 SPI1_SCK all AF5
ldr r1,[r0,#GPIOA+AFRL]; bic r1,r1,#(0x0f<<16); orr r1,r1,#(0x5<<16); str r1,[r0,#GPIOA+AFRL]
ldr r1,[r0,#GPIOA+AFRL]; bic r1,r1,#(0x0f<<24); orr r1,r1,#(0x5<<24); str r1,[r0,#GPIOA+AFRL]
ldr r1,[r0,#GPIOA+AFRL]; bic r1,r1,#(0x0f<<28); orr r1,r1,#(0x5<<28); str r1,[r0,#GPIOA+AFRL]
ldr r1,[r3,#AFRH]; bic r1,r1,#(0x0f << 12); orr r1,r1,#(0x5<<12); str r1,[r3,#AFRH]
@configure PA8 and PC9, PA4, PA6, PA7 and PG11 for alternates
ldr r1,[r0,#GPIOA+MODER]; bic r1,r1,#(0x3<<16); orr r1,r1,#(0x2<<16); str r1,[r0,#GPIOA+MODER]
ldr r1,[r0,#GPIOA+MODER]; bic r1,r1,#(0x3<<8); orr r1,r1,#(0x2<<8); str r1,[r0,#GPIOA+MODER]
ldr r1,[r0,#GPIOA+MODER]; bic r1,r1,#(0x3<<12); orr r1,r1,#(0x2<<12); str r1,[r0,#GPIOA+MODER]
ldr r1,[r0,#GPIOA+MODER]; bic r1,r1,#(0x3<<14); orr r1,r1,#(0x2<<14); str r1,[r0,#GPIOA+MODER]
ldr r1,[r0,#GPIOC+MODER]; bic r1,r1,#(0x3<<18); orr r1,r1,#(0x2<<18); str r1,[r0,#GPIOC+MODER]
ldr r1,[r3,#MODER]; bic r1,r1,#(0x3<<22); orr r1,r1,#(0x2<<22); str r1,[r3,#MODER]

@enable TIM2 to count down
ldr r0,=TIM2_BASE; ldr r1,=0x11; str r1,[r0,#TIM2_CR1]
ldr r0,=SPI1_BASE
@32-bit data /16 prescaler SPI1_CFG1
ldr r1,=(0x03<<28)|(0x1f<<0); str r1,[r0,#SPI1_CFG1]
@master Hardware NSS SPI1_CFG2
ldr r1,=(1<<31)|(1<<29)|(1<<22); str r1,[r0,#SPI1_CFG2]
@TSIZE set to 1 SPI1_CR2
ldr r1,=0x00; str r1,[r0,#SPI1_CR2]
@enable SPI peripheral SPE bit 0 in SPI1_CR1
ldr r1,[r0,#SPI1_CR1]; orr r1,r1,#0x1; str r1,[r0,#SPI1_CR1]
@start transfer CSTART bit 9 in SPI1_CR1
ldr r1,[r0,#SPI1_CR1]; orr r1,r1,#(1<<9); str r1,[r0,#SPI1_CR1]
@reload TXDR and then delay
ldr r1,=TIM2_BASE; ldr r3,[r1,#TIM2_CNT]
tx_loop: str r3,[r0,#SPI1_TXDR]; subs r3,r3,#0x1000
time1: ldr r4,[r1,#TIM2_CNT]; subs r4,r4,r3; bpl time1; b tx_loop

.ltorg; .org 0x200

 

Best answer by rich.g.williams

Hello,

Thank you for asking. I would say the main motivation is learning and understanding the device, rather than optimisation.

I recognise that the STM32H5E5 is an astonishingly capable device, and I certainly would not suggest that assembler is required to make good use of it. In fact, I don't yet have an application that really justifies all of its capabilities.

Part of the reason for using assembler is simply that it plays to my strengths. I am much more comfortable working close to the hardware and thinking in terms of registers, memory and individual processor operations than I am working in C. I know that C is the appropriate and productive choice for many STM32 applications, but I find that starting at the lowest practical level helps me understand the hardware before introducing another layer of abstraction.

That has also been particularly useful for the H5E5 boot process. Working through reset, the system bootloader, RAM execution, vector tables, security configuration, GPIO and now SPI has allowed me to establish what the hardware is actually doing rather than simply assuming that the generated software is doing it correctly.

There is another reason for doing it this way. I am posting the work on the forum because I hope that the experiments and the resulting minimal code may provide some visibility for other engineers who are approaching the H5E5, or similar devices, and want to understand what is happening underneath the higher-level software.

So there is certainly an element of learning, and an element of enjoyment as well. The small amount of code is really a consequence of the approach rather than the objective.

I have a great deal of respect for what ST has put into the H5E5. My intention is not to replace the normal STM32 development approach, but to understand what is underneath it and hopefully make some of that understanding useful to others.

6 replies

Associate II
September 23, 2026

Thanks for your sharing. Looks great. Do you have any more comments about the test result?

Associate II
September 24, 2026

Yes and thanks, here are the comments as a summary description of what the code does, any comments you may have much appreciated:-

 

===========================================================================
  SUMMARY: BARE-METAL CLOCK, VOS0, MCO AND SPI1 TEST MODULE
===========================================================================

1. PURPOSE
--------------------------------------------------------------------------------
This is a small bare-metal ARM Thumb-2 test program executed directly from SRAM.
It is intended to verify the STM32H5E5 clock and power system in stages, without
HAL or other startup software.

The tests provide external and timing evidence that each stage has operated
correctly before the next stage is used.

The main objectives are:
  • Change the core voltage from the reset state VOS3 to VOS0.
  • Bring PLL1 and PLL2 online.
  • Run SYSCLK at 248 MHz.
  • Output selected PLL clocks on MCO1 and MCO2 for direct measurement.
  • Run SPI1 from the PLL2 clock and verify its operation externally.
  • Use TIM2 as an independent timing reference during the tests.

The program executes from SRAM, so Flash wait-state settings are deliberately
left untouched.  The same clock configuration will require appropriate Flash
latency if the program is later executed from Flash.


2. TEST SEQUENCE
--------------------------------------------------------------------------------

[STAGE 1: VOS3 -> VOS0]
  • Requests VOS0 directly from the reset state VOS3.
  • Polls VOSRDY until the requested voltage is ready.
  • This verifies that the core voltage transition completes before the
    high-frequency clock is applied.


[STAGE 2: PLL1 AND PLL2 / 248 MHz SYSCLK]
  • Configures and enables PLL1 for a 248 MHz PLL1P output.
  • Configures and enables PLL2 for a 64 MHz PLL2P output.
  • Waits for both PLLs to report ready.
  • Switches SYSCLK to PLL1P.
  • TIM2 is then used to confirm the resulting increase in processor/timer
    clock rate by comparing measured delays.


[STAGE 3: MCO CLOCK OUTPUT TEST]
  • Routes selected PLL outputs to MCO1 and MCO2.
  • MCO1 is output on PA8 and MCO2 on PC9.
  • The divided clock outputs can be measured directly with an oscilloscope
    or frequency counter.
  • This provides an independent check of PLL frequency and clock routing,
    rather than relying only on register settings.


[STAGE 4: SPI1 TEST]
  • Enables the SPI1 peripheral clock and configures its GPIO alternate
    functions.
  • SPI1 is configured as a master using the PLL2P clock.
  • The SPI clock is divided to give the required test data rate.
  • 32-bit test values are written directly to SPI1_TXDR.
  • TIM2 supplies the timing reference for repeated transmissions.
  • The SPI1 clock and data can be observed externally to verify the SPI
    clock rate, framing and continuous operation.


3. TEST PHILOSOPHY
--------------------------------------------------------------------------------
The program deliberately keeps the startup sequence small and explicit.
Each major clock or power change has a corresponding observable result:

  VOS0  ->  power/current change and VOSRDY
  PLL1  ->  248 MHz SYSCLK and MCO output
  PLL2  ->  64 MHz peripheral clock and MCO output
  TIM2  ->  measurable processor/timer timing
  SPI1  ->  externally observable clock and data

This makes the module useful as a hardware bring-up and verification program
before the code is developed into a more complete startup sequence.

Associate II
September 24, 2026

The code was developed and tested on the ST NUCLEO-H5E5ZJ PCB. During the learning and testing process I drafted a compacted PCB schematic in two parts.  You can find the original schematics on this ST Forum thread:-  

NUCLEO-H5E5ZJ — Redrawing the Schematics as a GPIO View

For this startup test code I have drafted an additional skin schematic showing the SPI1 signals and the MCO1 and MCO2 outputs as configured on the NUCLEO-H5E5ZJ PCB.

That said the startup test code will likely run on any custom STM32H5E5 and STM32H5F5 PCB.

Here is the skin schematic attached 

mƎALLEm
ST Technical Moderator
September 25, 2026

Hello ​@rich.g.williams ,

Thank you for the effort. I’m curious to know about the motivation behind the coding with assembler. Is it for learning purpose or optimization purposes or just coding for fun?  because you already know STM32H5E5ZJ is a 4MB device running up to 250MHz.

To give better visibility on the answered topics, please click "Best answer" on the reply which solved your issue or answered your question.
rich.g.williamsAuthorBest answer
Associate II
September 25, 2026

Hello,

Thank you for asking. I would say the main motivation is learning and understanding the device, rather than optimisation.

I recognise that the STM32H5E5 is an astonishingly capable device, and I certainly would not suggest that assembler is required to make good use of it. In fact, I don't yet have an application that really justifies all of its capabilities.

Part of the reason for using assembler is simply that it plays to my strengths. I am much more comfortable working close to the hardware and thinking in terms of registers, memory and individual processor operations than I am working in C. I know that C is the appropriate and productive choice for many STM32 applications, but I find that starting at the lowest practical level helps me understand the hardware before introducing another layer of abstraction.

That has also been particularly useful for the H5E5 boot process. Working through reset, the system bootloader, RAM execution, vector tables, security configuration, GPIO and now SPI has allowed me to establish what the hardware is actually doing rather than simply assuming that the generated software is doing it correctly.

There is another reason for doing it this way. I am posting the work on the forum because I hope that the experiments and the resulting minimal code may provide some visibility for other engineers who are approaching the H5E5, or similar devices, and want to understand what is happening underneath the higher-level software.

So there is certainly an element of learning, and an element of enjoyment as well. The small amount of code is really a consequence of the approach rather than the objective.

I have a great deal of respect for what ST has put into the H5E5. My intention is not to replace the normal STM32 development approach, but to understand what is underneath it and hopefully make some of that understanding useful to others.

mƎALLEm
ST Technical Moderator
September 25, 2026

Thank you for these inputs. I need to modify the title to be inline with the content mainly the assembler wording that needs to be exposed in the title and let me mark your answer as Best Answer to better visibility for other users.

To give better visibility on the answered topics, please click "Best answer" on the reply which solved your issue or answered your question.