Skip to content

Latest commit

 

History

History

Folders and files

NameName
Last commit message
Last commit date

parent directory

..
 
 
 
 
 
 

README.md

Optimization Targets Sample

This sample is an FPGA tutorial that demonstrates how to set optimization targets for your compile to target different performance metrics.

This tutorial shows compiling with the minimum latency optimization target to achieve low latency at the cost of reduced fMAX.

Area Description
What you will learn How to set optimization targets for your compile.
How to use the minimum latency optimization target to compile low-latency designs.
How to manually override underlying controls set by the minimum latency optimization target.
Time to complete 20 minutes
Category Concepts and Functionality

Purpose

This FPGA tutorial demonstrates how to set optimization targets for your compile to target different performance metrics.

The -Xsoptimize=<flag> command-line option sets optimization targets, and it supports the following flags:

Flag Explanation Documentation
latency Minimum latency: minimize kernel latency at the cost of decreased fMAX Minimum Latency Flow
throughput-area-balanced Balanced throughput-area trade-offs: disable throughput-area trade-off heuristics that increase the throughput at the cost of area Balanced Throughput-Area Trade-Offs Flow
area Minimum area: minimize kernel area at the cost of decreased fMAX Minimum Area Flow

To compile your design with the minimum latency optimization target, use the flag option -Xsoptimize=latency.

As an example, this tutorial shows how to use the minimum latency optimization target to compile low-latency designs and how to manually override underlying controls set by the minimum latency optimization target. By default, the minimum latency optimization target tries to achieve lower latency at the cost of decreased fMAX, so it is a good starting point for optimizing latency-sensitive designs.

Prerequisites

Optimized for Description
OS Ubuntu* 20.04, Ubuntu* 22.04, Ubuntu* 24.04
RHEL* 9
SUSE* 15
NOTE: Windows is not supported
Hardware Agilex® 3, Agilex® 5, Agilex® 7, Stratix® 10 and Arria® 10 FPGAs
Software HLS IP Gen Compiler

Note: Even though the HLS IP Gen compiler is enough to compile for emulation, generating reports and generating RTL, there are extra software requirements for the simulation flow and FPGA compiles.

For using the simulator flow, Quartus® Prime Pro Edition (or Standard Edition when targeting Cyclone® V) and one of the following simulators must be installed and accessible through your PATH:

  • Questa*-Altera® FPGA Edition
  • Questa*-Altera® FPGA Starter Edition
  • ModelSim® SE

When using the hardware compile flow, Quartus® Prime Pro Edition (or Standard Edition when targeting Cyclone® V) must be installed and accessible through your PATH.

Warning: Make sure you add the device files associated with the FPGA that you are targeting to your Quartus® Prime installation.

This sample is part of the FPGA code samples. It is categorized as a Tier 3 sample that demonstrates a compiler feature.

flowchart LR
   tier1("Tier 1: Get Started")
   tier2("Tier 2: Explore the Fundamentals")
   tier3("Tier 3: Explore the Advanced Techniques")
   tier4("Tier 4: Explore the Reference Designs")

   tier1 --> tier2 --> tier3 --> tier4

   style tier1 fill:#0071c1,stroke:#0071c1,stroke-width:1px,color:#fff
   style tier2 fill:#0071c1,stroke:#0071c1,stroke-width:1px,color:#fff
   style tier3 fill:#f96,stroke:#333,stroke-width:1px,color:#fff
   style tier4 fill:#0071c1,stroke:#0071c1,stroke-width:1px,color:#fff
Loading

Find more information about how to navigate this part of the code samples in the FPGA top-level README.md. You can also find more information about troubleshooting build errors, links to selected documentation, and more.

Key Implementation Details

The sample illustrates the following important concepts.

  • Setting optimization targets to use when compiling your program.
  • Using the minimum latency optimization target to compile low-latency designs.
  • Manually overriding underlying controls set by the minimum latency optimization target.

Understanding the Tutorial Design

The basic function performed by the tutorial kernel is an RGB to grayscale algorithm. We compile the design three times to see the impact of the minimum latency optimization target in this tutorial in terms of latency and fMAX and to see how to override underlying controls set by the minimum latency optimization target with specific manual controls.

  • Part 1 compiles the design without the -Xsoptimize=latency flag. In this default flow, the compiler targets higher throughput and fMAX with the sacrifice of latency and area.

  • Part 2 compiles the design with the -Xsoptimize=latency flag, so the minimum latency optimization target is used in this compile, which lowers latency by trading off fMAX.

  • Part 3 compiles the design with the minimum latency optimization target and includes manual controls that revert default underlying controls set by the minimum latency optimization target. Therefore, latency and fMAX of this compile are the same as part 1.

Build the Optimization Targets Tutorial

Note: When working with the command-line interface (CLI), you should configure the HLS IP Gen Compiler using environment variables. Set up your CLI environment by sourcing the fpgavars script in the root of your HLS IP Gen Compiler installation every time you open a new terminal window. This practice ensures that your compiler, libraries, and tools are ready for development.

Linux*:

  • source <install-dir>/fpgavars.sh
  • For non-POSIX shells, like csh, use the following command: bash -c 'source <install-dir>/fpgavars.sh ; exec csh'

On Linux*

  1. Change to the sample directory.

  2. Build the program for Agilex® 7 device family, which is the default.

    mkdir build
    cd build
    cmake .. -DPART=<X>
    

    where -DPART=<X> is:

    • -DPART=NO_CONTROL: Compile part 1
    • -DPART=MINIMUM_LATENCY: Compile part 2
    • -DPART=MANUAL_REVERT: Compile part 3

    Note: You can change the default target by using the following command. Targeting a BSP is not supported.

    cmake .. -DFPGA_DEVICE=<FPGA device family or FPGA part number>
    
  3. Compile the design. (The provided targets match the recommended development flow.)

    1. Compile and run for emulation (fast compile time, targets emulates an FPGA device).
      make fpga_emu
      
    2. Generate the HTML optimization reports. (See Read the Reports below for information on finding and understanding the reports.)
      make report
      
    3. Compile for simulation (fast compile time, targets simulated FPGA device).
      make fpga_sim
      
    4. Compile and run on FPGA hardware (longer compile time, targets an FPGA device).
      make fpga
      

Read the Reports

After compiling the three reports in different cmake folders, locate all the the report.html files:

  • Report-only compile: optimization_targets.report.prj
  • FPGA hardware compile: optimization_targets.fpga.prj

Navigate to Loop Analysis (Throughput Analysis > Loop Analysis). In this viewer, you can find the latency of loops in the kernel. The latency of the compile with the minimum latency optimization target (part 2) should be lower than the other two compiles. Also, the latency of the other two compiles (part 1 & 3) should be the same.

Navigate to Clock Frequency Summary (Summary > Clock Frequency Summary) in optimization_targets.fpga.prj/reports/report.html (after make fpga completes). In this table, you can find the actual fMAX. The fMAX of the compile with the minimum latency optimization target (part 2) should be lower than the other two compiles. Also, the fMAX of the other two compiles (part 1 & 3) should be the same. Note that only the report generated by the FPGA hardware compile will reflect the true fMAX affected by the minimum latency optimization target. The difference is not apparent in the reports generated by make report because a design's fMAX cannot be predicted.

Run the Optimization Targets Sample

On Linux

  1. Run the sample on the FPGA emulator (the kernel executes on the CPU).

    ./optimization_targets.fpga_emu
    
  2. Run the sample on the FPGA simulator device.

    CL_CONTEXT_MPSIM_DEVICE_INTELFPGA=1 ./optimization_targets.fpga_sim
    

Example Output

Example output without minimum latency optimization target (part 1):

Kernel Throughput: 195.716MB/s
Exec Time: 1.9491e-05s, InputMB: 0.0038147MB
PASSED: all kernel results are correct

Example output with minimum latency optimization target (part 2):

Kernel Throughput: 137.764MB/s
Exec Time: 2.769e-05s, InputMB: 0.0038147MB
PASSED: all kernel results are correct

Example output with minimum latency optimization target but controls manually reverted (part 3):

Kernel Throughput: 192.934MB/s
Exec Time: 1.9772e-05s, InputMB: 0.0038147MB
PASSED: all kernel results are correct

Comparing to Arria® 10 GX FPGA, it is more notable on Stratix® 10 SX FPGA that the minimum latency optimization target significantly reduces the latency, along with the fMAX and the throughput. That is because the minimum latency optimization target disables the hyper-optimized handshaking, which achieves higher fMAX at the cost of increased latency.

Note: For more information on the hyper-optimized handshaking protocol on Stratix® 10 and Agilex® 7 devices, see the Modify the Handshaking Protocol Between Clusters (-Xshyper-optimized-handshaking) topic in the HLS IP Gen Handbook.

License

Code samples are licensed under the MIT license. See License.txt for details.

Third-party program Licenses can be found here: third-party-programs.txt.