[LCRC Accounts] Project Request: orfpga
Hello, A new project on the LCRC cluster has been requested. Please forward the information on to the LCRC Allocation sub-committee. Applicant's name: Azamat Mametjanov Applicant's institution: ANL Applicant's division: MCS Project Name: orfpga Project title: Autotuning FPGA designs Associated funding: NASA Other Systems: Science: Designing FPGAs for space applications is a complex iterative process involving simultaneous optimization of timing score, device utilization, and power consumption cite{Gaisler1,Gaisler2}. Achieving near optimal performance is challenging even for the most experienced engineers. The number of vendors offering FPGAs (e.g., Xilinx, Altera, QuickLogic, Actel/Microsmei) for space applications has also increased. Each of these vendors base their FPGAs on custom architectures and thus, the best performing design for a given application varies depending on the chosen FPGA architecture. In addition, the optimizations parameters, techniques, and best practices often differ for each vendor's FPGA. For organizations that use FPGAs from multiple vendors, this exacerbates the already lengthly FPGA design optimization process. The FPGA design techniques have evolved to provide abstractions such as IPCores (Intellectual Property Cores), which provide individual functions that can be composed to build larger designs. Therefore, in order to reduce development effort and increase code reuse, IPCores are often used in FPGA designs. However, IPCores have many tunable parameters cite{Junchao} that affect the performance of an FPGA design. In addition, the designer wants to experiment with various HDL (VHDL, Verilog etc.) and HLS (high level synthesis) code constructs. For example, the user might want to try different pipeline depths for a given piece of HLS code. Lastly, the designer also needs to experiment with different input parameters and random seeds to the synthesis and implementation (Mapping, Placement and Routing (PAR)) tools. These user settings can have dramatic impact on design performance because they influence the usage of underlying optimizations such as merging, retiming, and pipelining cite{SE1}. Thus the user tunable parameter space includes: 1) user customizable IPCores parameters, 2) various HDL and HLS code constructs, and 3) input parameters to the Synthesis, Mapping, and Placement and Routing tools. The choice of these parameters and the underlying FPGA architecture significantly affects the performance of an FPGA design. Manually exploring the aforementioned FPGA parameter space is a time consuming task and impedes designer productivity. The existing FPGA design tools (e.g., Xilinx Vivado/ISE Design Suites, Altera Quartus II) are unable to automatically deliver the best possible design for an application through the underlying optimization algorithms. Further HDL/HLS code optimization, and tuning of parameters is required at the designer level to meet the performance goals. This situation results in a potentially large gap between the achieved FPGA performance and best achievable performance on a given FPGA architecture. There is a need among NASA FPGA developers for next generation tools that can generate performance models and assist in deducing the near optimal designs through higher level inputs (performance goals, possible parameter values, different code variants etc.). Such performance optimization tools cite{Vuduc,Hart2009Orio,OrCuda2012, WN147,Tiwari2011,BalWilNor2011} are very prominent in high performance computing (HPC) domain. For scientific computing codes, these autotuning tools greatly assist performance optimizations and improve developer productivity. For example, the Orio tool cite{Hart2009Orio, OrCuda2012} (developed by ANL) takes annotated C or FORTRAN source code as input, generates optimized code variants from the annotated code, and empirically evaluates the performance of the generated codes. Orio then selects the best performing version to use for production runs. Because the search space of all possible optimized code versions can be huge and brute-force search strategy may not always be feasible, Orio provides various search methods (e.g., random search, simulated annealing) for reducing the search space and thus the empirical testing time. Significant performance gains in real scientific applications (an order of magnitude speedups for plasma physics code and climate modeling applications cite{Chung}) have been realized using the autotuning tools. There are no similar autotuning tools for FPGA design. In order to apply similar principles to FPGA designs, additional research and development is necessary. Significant gaps must be filled to automate the steps required for creating performance models on which the FPGA design decisions can be made. Building FPGA designs (synthesis and implementation) suffers from long execution times resulting in enormous compute time for empirical based autotuning. To be suitable for use in NASA labs, it is necessary to investigate new methodologies for optimizing the turnaround time. New automation tools need to be developed targeting the annotation of HDL/HLS code and FPGA design tool scripts. Finally, it is also required to develop a set of robust easy-to-use tools that interface with the existing FPGA design tools. By filling these gaps, the autotuning techniques from the HPC community can be made desirable and beneficial to NASA FPGA developers and the wider hardware development community. Project description: RNET and ANL are proposing to fully develop an empirical performance optimization tool called OrFPGA that efficiently explores the ``user tunable'' parameter space of an FPGA design and assists in deducing the near optimal design in terms of timing score, device utilization, and power consumption. The tunable parameter space will include IPCore parameters, HDL and HLS code constructs, and parameter settings for the vendor's design tools. Special automation tools will be devleoped to facilitate annotation of HDL/HLS code and design tool scripts. The computational magnitude of empirical performance tuning of FPGA designs will be addressed by novel machine learning based search algorithms (requiring minimal empirical evaluations), computational steering, leveraging intermediate performance analysis results, and parallelization techniques. The tool will support specification of prioritized performance metrics, easy-to-use interfaces for defining the parameter space, and intuitive visualization of performance models. The user will be able to accurately identify the optimal power consumption by including the environmental settings in search space and through real-time power monitoring. The benefits of the tool to NASA will be demonstrated in terms of performance metrics and cost benefits (user productivity) using real NASA designs that are used in space missions. Industry partnership: RNet Technologies, Inc. Project URL: Requested allocation: 400000 Q1: 100000 Q2: 100000 Q3: 100000 Q4: 100000 Justification: Storage requirements: The requester has used undetermined amount hours of their initial startup project. In addition to approving an initial amount, please specify a Category and Subcategory for this project. For a list of the current selection of approved categories, please see: https://wiki.lcrc.anl.gov/wiki/Processes/Categories Once the Allocation committee has approved the project, please go to the Project Management page to create it: https://accounts.lcrc.anl.gov/projects.php Thank You, The LCRC Accounts System
participants (1)
-
accounts@lcrc.anl.gov