Running an MCU at high speeds can come with drawbacks depending on where the code is being executed from. Code located in flash memory will take more time to execute because flash memory has slower access time and consequently more wait states. The same code can be executed from a faster memory with less or no wait states such as RAM. It is therefore necessary to partition the code in such a way that performance is optimized based on application requirements. The SPARK Wireless Core is interrupt driven and should be optimized to execute as fast as possible to let the application have as much processing time as possible. In this section, we will describe some optimizations that can be done to improve the processing speed of the system.
Some MCUs implement instruction caches to improve the execution time of code stored in RAM. This greatly reduces the wait time for an instruction fetch once it is stored in cache. Our QUASAR EVK supports this feature using the STM32U5’s ICACHE module. This feature still requires the cache to be loaded first which will induce a higher processing time on the first call of some function.
In this section, we provide general guidelines that can be followed to optimize MCU performance. The examples shown are specific to the STM32G474 MCU series from STMicroelectronics use on our EVK1.4 board, but could be adapted to other MCU families as well. The reader should be familiar with the concepts of linker script, startup script and memory layout of the target MCU.
Some MCUs include two banks of memory: the standard SRAM and another bank of Core-Coupled-Memory (a.k.a. CCM RAM for STM devices or Fast RAM for others). The CCM RAM is faster than the standard SRAM and is usually smaller. We can optimize code execution time by moving the code from Flash memory to the CCM SRAM. You need to manually edit the linker script file to assign code to these specific memory locations.
Figure 103: Example memory layout (taken from the STM32G474 reference manual).
The MEMORY portion of the linker script tells the linker about the memory spaces available to it within the device. Two new entries (ISRCCMRAM, CCMRAM) have been added to define the CCMRAM region. The purpose of the first entry is to map the interrupt vector table at the beginning of the CCMRAM area. The second entry is used to map user functions.
Note
For clarity purposes, only the relevant code section are shown.
References to the start and end of these sections are needed as the code will need to be copied from flash memory to CCMRAM on boot. The following variables are used:
_isr_vector_start
_isr_vector_end
_ccmram_start
_ccmram_end
These can be found under the “section” delimiter of the linker script file.
SECTIONS{.../* The startup code into "FLASH" Rom type memory */.isr_vector:{.=ALIGN(4);_isr_vector_start=.;KEEP(*(.isr_vector))/* Startup code */.=ALIGN(4);_isr_vector_end=.;}>ISRCCMRAMAT>FLASH.ccmram:{.=ALIGN(4);_ccmram_start=.;/* create a global symbol at CCMRAM start */INCLUDEccmram_sections.ld/* user defined ccmram sections */.=ALIGN(4);_ccmram_end=.;/* define a global symbol at the end of the used CCMRAM section */}>CCMRAMAT>FLASH...}
Note that the line “INCLUDE ccmram_sections.ld” refers to a file where the user can add function to be stored to CCMRAM. See Adding User Functions to CCMRAM.
The code is initially stored in flash memory, it needs to be copied from flash memory to CCMRAM where it will be executed. References to the flash memory addresses are needed. These are stored in the following variables:
_ccmram_loadaddr
_isr_vector_loadaddr
They can be found under the “section” delimiter of the linker.
SECTIONS{.../* Used by the startup script to copy user-defined content from the "ccmram" section load address, * which is located in flash, into the "_ccmram_start" relocation address which is located in the * actual CCMRAM. */_ccmram_loadaddr=LOADADDR(.ccmram);/* Used by the C program to copy the vector table from the "isr_vector" section load address, * which is located in flash, into the relocation address which is located in the actual CCMRAM. * This is optional as the relocation won't take effect until the SCB->VTOR register is adjusted. */_isr_vector_loadaddr=LOADADDR(.isr_vector);...}
On boot, the reset handler copies global variables and constants to SRAM so that they are accessed more easily at run time. User functions from “ccram_sections.ld ” are copied to CCMRAM.
.section .text.Reset_Handler
.weak Reset_Handler
.type Reset_Handler, %function
Reset_Handler:
ldr r0, =_estack
mov sp, r0 /* set stack pointer */
/* Copy the data segment initializers from flash to SRAM */
ldr r0, =_sdata
ldr r1, =_edata
ldr r2, =_sidata
movs r3, #0
b LoopCopyDataInit
CopyDataInit:
ldr r4, [r2, r3]
str r4, [r0, r3]
adds r3, r3, #4
LoopCopyDataInit:
adds r4, r0, r3
cmp r4, r1
bcc CopyDataInit
/* Copy user functions from flash to CCMRAM from ccram_sections.ld */
ldr r0, =_ccmram_start
ldr r1, =_ccmram_end
ldr r2, =_ccmram_loadaddr
movs r3, #0
b LoopCopyCCMRAMfunctions
CopyCCMRAMfunctions:
ldr r4, [r2, r3]
str r4, [r0, r3]
adds r3, r3, #4
LoopCopyCCMRAMfunctions:
adds r4, r0, r3
cmp r4, r1
bcc CopyCCMRAMfunctions
/* Zero fill the bss segment. */
ldr r2, =_sbss
ldr r4, =_ebss
movs r3, #0
b LoopFillZerobss
FillZerobss:
str r3, [r2]
adds r2, r2, #4
LoopFillZerobss:
cmp r2, r4
bcc FillZerobss
/* Call the clock system initialization function.*/
bl SystemInit
/* Call static constructors */
bl __libc_init_array
/* Call the application's entry point.*/
bl main
A system initialization function is called before running the main program (see previous code snippet). This SystemInit call is in fact a C function where the interrupt vector table is relocated based on the variables from the linker script. This is shown below.
voidSystemInit(void){uint32_t*ptr_isr_vector_start=&_isr_vector_start;uint32_t*ptr_isr_vector_loadaddr=&_isr_vector_loadaddr;/* FPU settings ------------------------------------------------------------*/#if (__FPU_PRESENT == 1) && (__FPU_USED == 1)SCB->CPACR|=((3UL<<(10*2))|(3UL<<(11*2)));/* set CP10 and CP11 Full Access */#endif/* Configure the Vector Table location add offset address -----------------*/#ifdef VECT_TAB_SRAMSCB->VTOR=SRAM_BASE|VECT_TAB_OFFSET;/* Vector Table Relocation in Internal SRAM */#elif VECT_TAB_FLASHSCB->VTOR=FLASH_BASE|VECT_TAB_OFFSET;/* Vector Table Relocation in Internal FLASH */#else/* Move the isr vector to CCMRAM-------------------------------------------*/while(ptr_isr_vector_start<&_isr_vector_end){*ptr_isr_vector_start++=*ptr_isr_vector_loadaddr++;}SCB->VTOR=CCMSRAM_BASE|VECT_TAB_OFFSET;/* Vector Table Relocation in CCMRAM */#endif}
When including a file in the linker script, it is mandatory to add the linker flag -L followed by the target file path. For more information, see the INCLUDE directive in the GNU LD Documentation’s Options Commands section.
Relocating functions to CCMRAM is done by specifying the input file, the section and/or the function name as described in the GNU LD Documentation’s Section Placement section.
Here are some examples:
Relocate the content of the text section from an SDK static library.
In the SDK release package, Wireless Core functions are provided through a prebuilt static library whose name depends on the selected build configuration. Use the actual library name available in core/wireless/prebuilt/ when writing linker-script placement rules.
The following pattern is intended to match any Wireless Core prebuilt library variant:
*swc_api_*_lib.a:(.text)
The leading * allows the rule to match the library even when the linker sees a path before the archive name. The second * matches the configuration-specific variant portion of the library name.
-O2 is the recommended optimization level for SDK applications. With the -O2 flag, GCC performs nearly all supported optimizations that do not involve a space-speed tradeoff. Compared to -O, this option increases both compilation time and the performance of the generated code.
Some optimizations might reduce the developer’s ability to debug the application. For example, some variables may be optimized out and may not be visible when debugging with GDB.
In the SDK release package, selected Wireless Core and SPARK Radio PHY implementation files are provided through prebuilt static libraries. These libraries are already compiled when the SDK release package is generated. Changing the application optimization level does not recompile the code contained in those prebuilt libraries.
The application optimization level still applies to the sources compiled as part of the user project, such as the application code, BSP, backend, middleware, modules, and other SDK source components that remain part of the release package.
Note
For best compatibility and predictability, application projects should use the SDK CMake configuration as the reference for compiler flags, compile definitions, include paths, linked libraries, and linker options.
The -fno-omit-frame-pointer flag is a compiler option in GCC that prevents the compiler from omitting the frame pointer in the generated assembly code. Preserving the frame pointer makes stack frames easier to interpret during debugging. It is common practice to use this flag in development builds or when debugging.
To optimize for performance and help with processing time, the flag can be removed. Note that this may have an effect on debugging.
In the SDK release package, removing or adding this flag only affects sources compiled by the application project. It does not change the code already compiled into the Wireless Core and SPARK Radio PHY prebuilt static libraries.
Important
Optimization flags may be changed for application sources, but CPU, FPU, floating-point ABI, and architecture flags must remain compatible with the prebuilt SDK static libraries. In particular, applications linking against the SDK release prebuilt libraries must use a compatible hard-float ABI configuration.
When using multiple optimization flags, the code execution order might change. This could cause issues in critical sections of the code. To prevent this, data synchronization instructions can be used to ensure that all memory accesses are completed before proceeding to the next instruction.
For example, in ARM Cortex-M based MCUs, the Data Synchronization Barrier (DSB) and Instruction Synchronization Barrier (ISB) instructions can be used to ensure that all memory accesses are completed before proceeding to the next instruction.
We recommend adding these instructions in the Board support package after context switch function calls or after enabling interrupts to ensure that the system is properly synchronized after a context switch has happened.
Most MCUs will have a feature in the SPI peripheral to add inter-byte spacing. This feature is required on the SR1000 series as noted in SPI Implementation. With the SR1100 series, the inter-byte spacing can be disabled to optimize the SPI throughput.