Projects

Research projects

A common thread runs through our projects: let the compiler and the microarchitecture share the work of extracting parallelism, and give each a program representation the other can act on directly.

Current projects

Vectorized Instruction Space (VIS)

Active
NSF CCF-2211353 · Collaborative Research: SHF: Medium · Oct 2022 – Sep 2027 · with Florida State University

How fast a processor can execute a program is still largely decided by how fast it can get through branch instructions. VIS applies the principle of vectorization to the instruction space: using an executable dynamic single-assignment form at the machine-instruction level, control dependences are converted into data dependences with far less overhead than traditional predication, naturally combining controlled multi-path execution with speculation. The goal is to eliminate a large fraction of hard-to-predict branches and to unroll loops with unknown trip counts efficiently. The project develops the foundations of the paradigm at both the microarchitecture and compiler level, targeting superscalar and VLIW machines, and releases its toolsets as open source.

Statically Controlled Asynchronous Lane Execution (SCALE)

Completed 2025
NSF CCF-1901005 · Collaborative Research: SHF: Medium · Sep 2019 – Aug 2025 · with Florida State University

SCALE aims to meet or exceed the performance of a traditional superscalar processor while approaching the energy efficiency of a VLIW. The architecture provides separate asynchronous execution lanes; dependences between instructions in different lanes are identified statically by the compiler, which inserts the inter-lane synchronization. Because the lanes are explicit, the compiler can generate code for different modes of execution — explicit packaging of parallel instructions, parallel and pipelined loop iterations, SPMD execution, and independent multi-threading — and adapt to whatever kind of parallelism is available at each point in a program. The project produced the Synchronized Lane Architecture (SLA), its code generator, and the tooling for bootstrapping a new ISA.

International Research Experiences with NTNU

Completed 2025
NSF OISE-2103105 · IRES Track I: Collaborative Research · Sep 2021 – Aug 2025 · with Florida State University and NTNU

An NSF International Research Experiences for Students (IRES) program that sent cohorts of Michigan Tech and Florida State doctoral students to the Computer Architecture Laboratory (CAL) at the Norwegian University of Science and Technology in Trondheim for ten-week research stays each summer. Working with the groups of Professors Magnus Själander and Magnus Jahre, students pursued compiler and architecture techniques for automatically improving the performance and energy efficiency of applications — the same theme that runs through our VIS and SCALE work. The collaboration with NTNU continues beyond the award.

FAST — Flexible Architecture Simulation Tool

Infrastructure
Long-running lab infrastructure · ADL / ADL++

FAST generates cycle-level microarchitecture simulators from descriptions written in our architecture description language (ADL). Both the instruction set and the microarchitecture are specified in ADL; a typical superscalar description runs 6,000–8,000 lines and yields a 30,000–35,000-line C++ simulator. A new simulator takes one to three person-months to produce instead of the twelve to eighteen typical of hand-coded simulators, which is what makes thorough design-space exploration practical. FAST includes a debugging environment and extensive performance-data collection, and version 7.2 can checkpoint and restart at any point — so every SimPoint of a benchmark can be simulated in parallel.

Earlier projects

  • Dependent ILP: Dynamic Hoisting and Eager Scheduling of Dependent Instructions NSF CCF-1823398 · FoMR: Collaborative Research · 2018 · with Florida State University Techniques for hoisting and eagerly scheduling instructions that depend on one another, to expose parallelism that conventional out-of-order scheduling leaves on the table.
  • Sphinx: Combining Data and Instruction Level Parallelism through Demand-Driven Execution of Imperative Programs NSF CCF-1533828 · XPS: Full · 2015 · with Florida State University · preceded by EAGER award CCF-1450062 (2014) Von Neumann execution does not scale well because communication, synchronization, and naming are hard to scale. Sphinx developed an alternative: programs compiled from ordinary languages are executed in a demand-driven fashion, harvesting both instruction- and data-level parallelism through compiler–architecture collaboration. The key is the Future Gated Single Assignment (FGSA) form, which is at once a representation for the optimizing compiler and the instruction set of the demand-driven machine. Papers supported by Sphinx include Dynamic Memory Dependence Predication (ISCA 2018) and A Two-Phase Recovery Mechanism (ICS 2018).
  • Single Assignment Architecture / Single Assignment Compiler NSF CCF-1116551 · SHF: Small · 2011 Established the single-assignment compiler / single-assignment architecture pairing, including Future Gated Single Assignment form and static single assignment with congruence classes (CGO 2014).
  • Future Values: Reshaping the Future of Instruction Level Parallelism NSF CCF-0347592 · CAREER · 2004 Program representations and processor mechanisms built on future values — referring to a value before the instruction that produces it — enabling unrestricted code motion and new ways to order instructions (U.S. Patent 7,747,993).
  • Exposing the Compiler to the Hardware: Memory Subsystem Optimizations through Compiler/Micro-architecture Cooperation using Set Membership Information and Color Sets NSF CCF-0312892 · ITR · 2003 · with Steven Carr Compiler-directed memory subsystem optimization: reuse-distance and store-distance analysis, feedback-directed memory disambiguation, and color-set-based memory dependence prediction.
  • Power-Adaptive Microarchitecture and Compiler Design for Mobile Computing DARPA Power Aware Computing/Communications Program · F29601-00-1-0183 · 2000 – 2002 Compiler and microarchitecture techniques for reducing static and dynamic power in superscalar processors, including optimizing static power dissipation by functional units.