<?xml version="1.0" encoding="UTF-8"?>
<?xml-stylesheet type="text/xsl" href="https://feeds.feedblitz.com/feedblitz_rss.xslt"?>
<rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	 xmlns:feedburner="http://rssnamespace.org/feedburner/ext/1.0">
<channel>
	<title>Computer Architecture Today</title>
	<atom:link href="https://www.sigarch.org/feed/" rel="self" type="application/rss+xml" />
	<link>https://www.sigarch.org</link>
	<description>Informing the broad computing community about current activities, advances and future directions in computer architecture.</description>
	<lastBuildDate>Tue, 14 Jul 2026 00:40:33 -0400</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>hourly</sy:updatePeriod>
	<sy:updateFrequency>1</sy:updateFrequency>
	
<image>
	<url>https://www.sigarch.org/wp-content/uploads/2017/03/logo_rgb.png</url>
	<title>Computer Architecture Today</title>
	<link>https://www.sigarch.org</link>
</image> 
<site xmlns="com-wordpress:feed-additions:1">125883397</site>
<meta xmlns="http://www.w3.org/1999/xhtml" name="robots" content="noindex" />
<item>
<feedburner:origLink>https://www.sigarch.org/when-ai-enters-the-architecture-design-loop-what-counts-as-a-contribution/</feedburner:origLink>
		<title>When AI Enters the Architecture Design Loop, What Counts as a Contribution?</title>
		<link>https://feeds.feedblitz.com/~/960306302/0/sigarch-cat/</link>
		<comments>https://feeds.feedblitz.com/~/960306302/0/sigarch-cat/#respond</comments>
		<pubDate>Tue, 14 Jul 2026 00:40:33 +0000</pubDate>
		<dc:creator><![CDATA[Vijay Janapa Reddi]]></dc:creator>
		<category><![CDATA[ACM SIGARCH]]></category>
		<category><![CDATA[Benchmarks]]></category>
		<category><![CDATA[Machine Learning]]></category>
		<category><![CDATA[Measurements]]></category>
		<category><![CDATA[Methodology]]></category>
		<guid isPermaLink="false">https://www.sigarch.org/?p=108942</guid>
		<description><![CDATA[<div><img width="300" xheight="123" src="https://www.sigarch.org/wp-content/uploads/2026/07/Architecture-2.0-loop-diagram-300x123.png" class="attachment-medium size-medium wp-post-image" alt="" style="max-width:100% !important;height:auto !important;margin-bottom:15px;margin-left:15px;float:right;"  /></div>AI is starting to shape architectural mechanisms, workloads, and evaluation. To make sense of it, we need a compact, shared way to preserve enough of that process for other groups to evaluate and build on AI-assisted claims. At the 53rd ISCA in Raleigh, AI for architecture stopped feeling like a side conversation. In the hallways, [&#8230;]]]>
</description>
				<content:encoded><![CDATA[<div><img width="300" height="123" src="https://www.sigarch.org/wp-content/uploads/2026/07/Architecture-2.0-loop-diagram-300x123.png" class="attachment-medium size-medium wp-post-image" alt="" style="margin-bottom:15px;margin-left:15px;float:right;" decoding="async" loading="lazy" /></div><h1 id="whenaientersthearchitecturedesignloopwhatcountsasevidence"><em style="color: #666666; font-size: 14px;">AI is starting to shape architectural mechanisms, workloads, and evaluation. To make sense of it, we need a compact, shared way to preserve enough of that process for other groups to evaluate and build on AI-assisted claims.</em></h1>
<p>At the 53rd ISCA in Raleigh, AI for architecture stopped feeling like a side conversation. In the hallways, the talk kept coming back to one thing. AI is starting to enter the architecture design loop, the repeated process of framing a problem, proposing or editing a design, measuring it, rejecting weak candidates, and deciding what to try next.</p>
<p>In particular, there were two deep-dive workshops coupled with other activities. The <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://sites.google.com/view/mlarchsys">MLArchSys</a> workshop added A³, a segment on agentic approaches to architecture, and the full-day <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://harvard-edge.github.io/isca-26-arch-2-workshop/">Architecture 2.0</a> workshop focused entirely on agentic design. Both drew well over a hundred people and were standing-room-only by the end. The same shift was visible in the main program, in a plenary panel on <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://iscaconf.org/isca2026/program/">research and education in the GenAI era</a>.</p>
<p>&nbsp;</p>
<div id="attachment_108943" style="width: 1034px" class="wp-caption aligncenter"><img fetchpriority="high" decoding="async" aria-describedby="caption-attachment-108943" class="wp-image-108943" src="https://www.sigarch.org/wp-content/uploads/2026/07/IMG_9857-3840x2880.jpg" alt="" width="1024" height="768" srcset="https://www.sigarch.org/wp-content/uploads/2026/07/IMG_9857-980x735.jpg 980w, https://www.sigarch.org/wp-content/uploads/2026/07/IMG_9857-480x360.jpg 480w" sizes="(min-width: 0px) and (max-width: 480px) 480px, (min-width: 481px) and (max-width: 980px) 980px, (min-width: 981px) 1024px, 100vw" /><p id="caption-attachment-108943" class="wp-caption-text">The Architecture 2.0 workshops at ISCA 2026.</p></div>
<div id="attachment_108944" style="width: 1034px" class="wp-caption aligncenter"><img loading="lazy" decoding="async" aria-describedby="caption-attachment-108944" class="wp-image-108944" src="https://www.sigarch.org/wp-content/uploads/2026/07/1783067364621.jpeg" alt="" width="1024" height="768" srcset="https://www.sigarch.org/wp-content/uploads/2026/07/1783067364621-980x735.jpeg 980w, https://www.sigarch.org/wp-content/uploads/2026/07/1783067364621-480x360.jpeg 480w" sizes="(min-width: 0px) and (max-width: 480px) 480px, (min-width: 481px) and (max-width: 980px) 980px, (min-width: 981px) 1024px, 100vw" /><p id="caption-attachment-108944" class="wp-caption-text">The MLArchSys workshops at ISCA 2026.</p></div>
<p>&nbsp;</p>
<p>Full rooms on a particular subject matter are undoubtedly a sign of community momentum. The opportunity now is to turn that momentum into a durable engineering practice. That will take shared evidence, reusable tools, and enough agreement for a claim to leave the room where it was born and still be checked, compared, taught, improved, or rejected by someone else. AI is already producing architectural ideas. The question is what must travel with those ideas for them to become engineering knowledge.</p>
<p>Suppose a paper reports an AI-generated memory prefetcher with a 15 percent speedup. The code runs, and the speedup reproduces under the reported setup. But the agent saw some workloads and not others, adapted to simulator feedback, tried many candidates, and picked this one. What, exactly, is the contribution here? The final prefetcher? The speedup? The prompt? The agent? Or the process that connected them?</p>
<p>Once AI chooses workloads, responds to feedback, and selects which candidate to report, the same uncertainty about what counts as the contribution reappears for every such result. As AI gains more influence, we have to decide which evidence should accompany a result, so that another group can tell which part actually holds. This blog post is about the evidence that should accompany them if they are to become engineering knowledge.</p>
<h2 id="thescaleoftheshift">The Scale of the Shift</h2>
<p>The workshops reflect a broader rise in AI-mediated systems research. A recent cross-stack survey of more than 7,800 arXiv papers found that the annual AI-for-systems publication volume grew roughly 23× from 2017 to 2025, and even faster in hardware and chip design (<a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://arxiv.org/abs/2602.15241">GenAI for Systems</a>). These categories reach beyond architecture, and not every paper in them runs an adaptive design loop. But where AI adapts to workloads, simulator feedback, or selection criteria, the final artifact can obscure how the result emerged. As this body of work grows, leaving that process implicit makes it harder to compare results or carry a finding from one group to the next.</p>
<div id="attachment_108945" style="width: 753px" class="wp-caption aligncenter"><img loading="lazy" decoding="async" aria-describedby="caption-attachment-108945" class="wp-image-108945 size-full" src="https://www.sigarch.org/wp-content/uploads/2026/07/growth.jpg" alt="" width="743" height="372" srcset="https://www.sigarch.org/wp-content/uploads/2026/07/growth.jpg 743w, https://www.sigarch.org/wp-content/uploads/2026/07/growth-480x240.jpg 480w" sizes="(min-width: 0px) and (max-width: 480px) 480px, (min-width: 481px) 743px, 100vw" /><p id="caption-attachment-108945" class="wp-caption-text"><strong>Figure 1:</strong> Annual AI-for-systems publication volume grew about 23× from 2017 to 2025 (a). The hardware and chip-design categories grew roughly 43× and 60×, respectively, compared with 21× for software (b).</p></div>
<p>Recent SIGARCH posts show the field working this out in public from different angles. <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://www.sigarch.org/computer-architectures-alphazero-moment-is-here/">Karu Sankaralingam</a> asks whether architecture has reached an AlphaZero moment, with evaluation, not idea generation, as the real bottleneck. <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://www.sigarch.org/architecture-systems-are-changing-the-architects-role-in-the-era-of-agentic-co-design/">Dimitrios Skarlatos</a> argues that agentic co-design is already reshaping the architect’s role and the hardware-software contract. <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://www.sigarch.org/how-ai-will-reshape-computer-systems-by-2035-a-jeffersonian-dinner-in-san-francisco-about-our-10000x-future/">Jeff Dean and David Patterson</a> project a 10,000× future built on compounding gains, one of them AI automating hardware design itself. Together, their arguments point toward a common question about what should count as evidence when AI helps produce a design. Answering it requires being precise about what changes when AI moves from a bounded tool to an actor in the design process.</p>
<h2 id="whatchangeswithagenticdesign">What Changes With Agentic Design</h2>
<p>AI for architecture means using learned or agentic systems to help shape architectural designs and the evidence used to evaluate them, rather than building hardware optimized to run AI workloads. This direction, framed in recent work on the <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://ieeexplore.ieee.org/document/10857820">foundations of AI agents for modern computer system design</a>, overlaps with software generation and electronic design automation (EDA), but it is distinct from both. An open-ended architecture agent can influence mechanisms, abstractions, workloads, simulator configurations, and interfaces whose effects propagate through many downstream programs and tools. That reach is what makes both its designs and its decisions worth scrutinizing.</p>
<p>Machine learning has entered architecture before. <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://ieeexplore.ieee.org/document/903263">Perceptron branch predictors</a>, <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://ieeexplore.ieee.org/document/4556714/">reinforcement-learning memory controllers</a>, <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://dl.acm.org/doi/10.1145/3466752.3480114">learned prefetchers such as Pythia</a>, surrogate models that approximate expensive simulations, and autotuners that automatically search configuration choices all used statistical learning to sharpen a mechanism or search a design space. In much of that work, ML was part of the artifact or a bounded optimizer, while the workloads, evaluator, and rules governing the search were set outside the model. The agentic shift is not a clean break from autotuning. It expands the scope and authority of the adaptive process. When a system can propose or edit mechanisms, call tools, choose workloads, adapt to simulator feedback, and influence which candidate survives, ML is no longer only inside the design. It starts to shape the claim we make about the design.</p>
<p>This concern predates AI. Human researchers explore design spaces, tune systems, and discard candidates, too, and research has always run on authors disclosing what others need to judge the work, backed by a degree of trust. What changes with an agent is how we scale. An adaptive system can make and revise these choices at machine speed across mechanism code, simulator configurations, workloads, tool calls, and selection criteria, often in response to the same evaluator that later supports the claim. The issue is not that a choice made by an AI system is inherently less trustworthy. It is that a large, tool-mediated search collapses into a final mechanism and a score, and the path that produced them disappears unless someone deliberately records it. The end goal is not to eliminate trust, but to keep that part of the methodology visible enough for others to assess the claim.</p>
<p>Gupta and colleagues’ <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://arxiv.org/abs/2602.22425">ArchAgent</a> makes this concrete by designing and implementing cache-replacement policies, not just their parameters. Starting from Mockingjay, a prior state-of-the-art policy, ArchAgent generated Policy31 for the single-core SPEC CPU 2006 suite, with a usage-intensity mechanism that its authors could inspect and test feature by feature. It also generated Policy12, which appeared to beat Mockingjay through what the paper calls a simulator escape, a higher score won by exploiting the simulator rather than the architecture. In ChampSim, unsupported bypassing of last-level cache writes was protected only by an assertion that optimized builds removed, so Policy12 looked faster because the bypassed writes vanished rather than being handled correctly.</p>
<p>The same agentic process produced both a genuine mechanism and a broken measurement, and the reported scores alone would not tell a reviewer which was which. The authors, to their full credit, caught the escape through manual inspection and reported it, exactly the kind of evidence future studies should preserve. A fuller record would not have found the bug automatically, but preserving the build configuration, the rejected policy, and the check that disqualified it would let others see why Policy12 failed and reuse that check in the next study.</p>
<div id="attachment_108969" style="width: 1090px" class="wp-caption aligncenter"><img loading="lazy" decoding="async" aria-describedby="caption-attachment-108969" class="wp-image-108969 size-full" src="https://www.sigarch.org/wp-content/uploads/2026/07/Artifact-only-review-versus-loop-aware-review.png" alt="" width="1080" height="520" srcset="https://www.sigarch.org/wp-content/uploads/2026/07/Artifact-only-review-versus-loop-aware-review.png 1080w, https://www.sigarch.org/wp-content/uploads/2026/07/Artifact-only-review-versus-loop-aware-review-980x472.png 980w, https://www.sigarch.org/wp-content/uploads/2026/07/Artifact-only-review-versus-loop-aware-review-480x231.png 480w" sizes="(min-width: 0px) and (max-width: 480px) 480px, (min-width: 481px) and (max-width: 980px) 980px, (min-width: 981px) 1080px, 100vw" /><p id="caption-attachment-108969" class="wp-caption-text"><strong>Figure 2:</strong> (Left) Artifact-only review sees a generated design and a reported number. (Right) Design loop-aware review keeps the artifact in view while adding the declared bounds, the evidence and failures, and a record of who could accept or reject the candidate.</p></div>
<p>ArchAgent also helps show where <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://www.sigarch.org/computer-architectures-alphazero-moment-is-here/">Karu Sankaralingam’s AlphaZero comparison</a> holds and where architecture departs from it. <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://arxiv.org/pdf/1712.01815">AlphaZero</a> discovered powerful Go strategies through self-play, but the board and the rules stayed fixed. Only the strategy could change. Depending on its permissions, an architecture agent can influence the strategy, the board, the rules, and the score used to judge it. Benchmarks can become data the agent adapts to, simulators can become environments it acts on through tool calls, metrics can become optimization targets, and interfaces define what actions it can take. That is why the claim must carry a record of the &#8220;design loop,&#8221; not just the artifact that emerged from it.</p>
<h2 id="whatevidencetopreserve">What Evidence to Preserve</h2>
<p>Architecture papers already describe mechanisms, baselines, workloads, simulators, and evaluation procedures, so this is not a call for longer methods sections. What they rarely preserve is how the search reached the reported design. A compact record would make a few things visible:</p>
<ul>
<li><strong>Bounds:</strong> what the agent could see and change, and what stayed fixed</li>
<li><strong>Feedback:</strong> how much simulator feedback it drew on, and how it steered the search</li>
<li><strong>Evidence and failures:</strong> what supported the reported result, and which candidates were rejected and why</li>
<li><strong>The decision:</strong> who could reject a candidate, who made the final call, and what would overturn the result</li>
</ul>
<p>None of it is exotic. It is the part of the process that an adaptive search tends to erase.</p>
<p>Machine learning has already faced a version of this gap. A released model or dataset often lacked sufficient context to assess its intended use, evaluation, or provenance. In response, <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://arxiv.org/abs/1810.03993">model cards</a> and <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://arxiv.org/abs/1803.09010">datasheets for datasets</a> provided the field with compact records that accompany the artifact, short enough to read yet specific enough to state what the work does and does not support. Architecture needs the same kind of record, extended from a finished artifact to the search that produced it.</p>
<p>The reporting burden should scale with how much authority the agent had. If AI only helped implement a mechanism specified by a human, ordinary artifact disclosure is probably enough. If it chose workloads, edited the design, adapted to simulator feedback, or determined which candidate was reported, some account of that process should accompany the result. A simple test is whether AI materially shaped the mechanism, workload, evaluator, stopping rule, rejection rule, or reported result. That record need not become a universal checklist, expose the model’s private reasoning, or promise an exact replay of a randomized search. Its format should emerge through use and revision rather than being settled in advance. What matters is how much of an AI-shaped process must stay visible for another group to see why a result survived and whether it holds under different assumptions.</p>
<p>Existing practice offers only partial precedents. Declaring the setup before a search resembles <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://www.cos.io/initiatives/prereg">preregistration</a>, keeping the evidence trail resembles <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://ctuning.org/ae/">artifact evaluation</a>, and preserving failed alternatives resembles ablation studies, which test the effect of changing one part of a design, as well as negative result reporting. The individual practices are not new. The change is that a single adaptive system can operate continuously across the design, workload, evaluator, and stopping rule, which are usually documented separately.</p>
<p>A record like this is a good-faith disclosure, not proof. An author can omit an inconveniently rejected candidate, and a reviewer cannot rerun an adaptive search to catch the omission, especially when the agent relies on a proprietary model that shifts over time and never repeats a run exactly. The record cannot stand on its own. What keeps it honest is disclosure scaled to the agent’s authority, read by reviewers rather than filed as a badge, and confirmed against evidence the search did not produce.</p>
<p>A result selected through adaptive evaluation should face at least one confirmation check outside the search, using held-out workloads, a second simulator, or a targeted test of the claimed mechanism. If the same agent tunes against the simulator that scores it, selects its evaluation workloads, and stops once the metric looks good, the evaluator has become part of the optimization loop, the architecture equivalent of training on the test set, or evaluation leakage. The check must also use measurements appropriate to the claim, and as a recent <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://www.sigarch.org/the-return-of-rigorous-full-system-timing-simulation/">SIGARCH post on full-system timing simulation</a> argues, simulation speed and fidelity are already in tension before an agent begins optimizing. Agent feedback makes the measurement window and metric part of the search surface, so authors should explain why the confirmation is credible and what evidence would overturn the result.</p>
<p>The same shift that put agents into the design loop is now putting them into the review loop. In systems research, agents already <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://arxiv.org/abs/2510.06189">drive the design loop</a> end-to-end, and elsewhere they draft and review their own papers, with a language model serving as the <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://arxiv.org/abs/2306.05685">judge</a>. An automated judge can share the blind spots of the system it reviews, so it does not replace the independent check. But it does raise the value of a record built to be machine-readable as well as human-readable, one that the next agent can use to rebuild the setup, rerun the disqualifying check, and test the claim rather than take a summary on faith.</p>
<h2 id="asharedlayerfortheloop">A Shared Layer for the Loop</h2>
<p>A record inside a single paper is a good start. It becomes a shared convention when authors use common fields and present supporting evidence in a form others can inspect. A reviewer can then challenge the record, a student can learn why the reported design survived, and another group can revisit a rejected candidate under the same conditions. Some variation in this scaffolding is healthy, but shared infrastructure gives groups a common base without requiring them to pursue the same research questions.</p>
<p>The computer architecture community has previously built shared responses to analogous coordination problems. <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://www.spec.org/">SPEC</a> provided common workloads, while simulators such as SimpleScalar and gem5 provided researchers with reusable experimental platforms. <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://mlcommons.org/benchmarks/">MLPerf</a> and the long line of prediction and prefetching championships showed that we can agree on workloads, rules, and scoreboards. These shared objects did not settle every question, but they gave the field durable things to run, dispute, teach, and improve. Benchmarks do not capture the path through an adaptive search, but they show how common boundaries make comparisons meaningful. Agentic design now needs a similar layer for search state, allowed actions, failures, and independent checks.</p>
<p>When Amir Yazdanbakhsh and I first articulated the Architecture 2.0 vision in a SIGARCH blog post in 2023, it was conceived <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://www.sigarch.org/architecture-2-0-why-computer-architects-need-a-data-centric-ai-gymnasium/">as a data-centric AI gymnasium</a>, a shared ecosystem of data, benchmarks, and tools for ML-assisted architecture research. Three years later, many of those building blocks have emerged, including knowledge benchmarks such as <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://quarch.ai/">QuArch</a>, assembled with the help of <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://arxiv.org/pdf/2510.22087">more than 140 contributors across 40 institutions</a>, capability evaluations such as <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://arxiv.org/abs/2607.03601">ArchEval</a>, and design-space infrastructure such as <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://dl.acm.org/doi/pdf/10.1145/3579371.3589049">ArchGym</a>. Building them has taught us something more important. The remaining challenge is not simply another benchmark, evaluation, or piece of infrastructure. A benchmark tests what a model knows, an evaluation tests what an agent can do, and infrastructure runs the search. Some of these systems log a run in detail, but that record stays inside the tool. What no published result yet carries with it is a portable account of how a study bounded its search, rejected candidates, and chose what to report, and that is the part we cannot supply by building one more tool.</p>
<h2 id="makingitroutine">Making It Routine</h2>
<p>Making this kind of evidence part of everyday research practice will take deliberate community effort. Machine-learning communities have shown one path. Benchmarks and competitions provide participants with shared tasks and rules, while model cards and datasheets establish shared reporting expectations. <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://neurips.cc/Conferences/2026/EvaluationsDatasetsHosting">NeurIPS requires submissions to include a paper checklist</a> addressing reproducibility, transparency, limitations, and experimental details, a step it <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://blog.neurips.cc/2021/03/26/introducing-the-neurips-2021-paper-checklist/">introduced</a> to help authors document the completeness and limits of their work. <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://icml.cc/Conferences/2023/PaperGuidelines">ICML has likewise published paper guidelines</a>, based on the NeurIPS checklist, that ask authors to document claims, limitations, code, data, and experimental details. A complementary perspective appears in the <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://www.jennwv.com/papers/realml.pdf">2022 FAccT paper by Smith and colleagues on REAL ML</a>, which argues that responsible machine learning depends not only on models and metrics, but also on documenting the broader research process.</p>
<p>Architecture conferences and workshops could experiment with a few concrete practices. Artifact-evaluation tracks can request versioned configurations, failures that affected the result, and at least one confirmation check that was not used to select the reported result. Competitions can specify workloads, allowed actions, limits on evaluator queries, stopping rules, and held-out tests to ensure scores remain comparable. We do not need to harden these practices into permanent rules at once, and venues can learn what helps reviewers and discard what does not. For proprietary work, the detailed record may remain internal, but a public claim still requires sufficient disclosure for outside groups to assess it. The simplest shared form for that disclosure is a single page. Authors could include or link to a one-page <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://arch2.mlsysbook.ai/book/appendices/appendix-b-design-loop-card/">design-loop card</a> summarizing the process in a consistent format.</p>
<div id="attachment_108946" style="width: 1090px" class="wp-caption aligncenter"><img loading="lazy" decoding="async" aria-describedby="caption-attachment-108946" class="wp-image-108946 size-full" src="https://www.sigarch.org/wp-content/uploads/2026/07/Loop-contract-lifecycle.png" alt="" width="1080" height="520" srcset="https://www.sigarch.org/wp-content/uploads/2026/07/Loop-contract-lifecycle.png 1080w, https://www.sigarch.org/wp-content/uploads/2026/07/Loop-contract-lifecycle-980x472.png 980w, https://www.sigarch.org/wp-content/uploads/2026/07/Loop-contract-lifecycle-480x231.png 480w" sizes="(min-width: 0px) and (max-width: 480px) 480px, (min-width: 481px) and (max-width: 980px) 980px, (min-width: 981px) 1080px, 100vw" /><p id="caption-attachment-108946" class="wp-caption-text"><strong>Figure 3:</strong> One possible one-page record makes the bounds, actions, feedback, evidence, failures, and final decision visible.</p></div>
<p>These conventions also need a public home alongside tools, benchmarks, failure cases, and examples. The <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://arch2.mlsysbook.ai/">Architecture 2.0 hub</a> is one possible starting point. Its value will depend on whether multiple groups use, challenge, revise, and help govern its contents.</p>
<p>No single convention will make AI for architecture an engineering discipline. The value of a shared record is that it lets a result move beyond the group that produced it so someone who was not there can check it, build on it, or challenge it when the evidence does not hold. The full rooms at ISCA were the momentum. Making the evidence travel with the design is what turns momentum into a discipline.</p>
<h2 id="abouttheauthor">About the Author</h2>
<p>Vijay Janapa Reddi is the Gordon McKay Professor of Electrical and Computer Engineering at Harvard University and a visiting professor at ETH Zurich. His work spans computer architecture, machine learning systems, and autonomous agents. He is Vice President and a board member of MLCommons and the author of the open-source <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://mlsysbook.ai/"><em>Machine Learning Systems</em></a> book.</p>
<p class="disclaim"><strong>Disclaimer:</strong> <em>These posts are written by individual contributors to share their thoughts on the Computer Architecture Today blog for the benefit of the community. Any views or opinions represented in this blog are personal, belong solely to the blog author and do not represent those of ACM SIGARCH or its parent organization, ACM.</em></p><Img align="left" border="0" height="1" width="1" alt="" style="border:0;float:left;margin:0;padding:0;width:1px!important;height:1px!important;" hspace="0" src="https://feeds.feedblitz.com/~/i/960306302/0/sigarch-cat">
]]>
</content:encoded>
			<wfw:commentRss>https://feeds.feedblitz.com/~/960306302/0/sigarch-cat/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
	<post-id xmlns="com-wordpress:feed-additions:1">108942</post-id></item>
<item>
<feedburner:origLink>https://www.sigarch.org/isca-2026-trip-report/</feedburner:origLink>
		<title>ISCA 2026 Trip Report</title>
		<link>https://feeds.feedblitz.com/~/960054881/0/sigarch-cat/</link>
		<comments>https://feeds.feedblitz.com/~/960054881/0/sigarch-cat/#respond</comments>
		<pubDate>Sat, 11 Jul 2026 01:51:03 +0000</pubDate>
		<dc:creator><![CDATA[Bingyao Li]]></dc:creator>
		<category><![CDATA[ACM SIGARCH]]></category>
		<category><![CDATA[ISCA]]></category>
		<category><![CDATA[Trip Report]]></category>
		<guid isPermaLink="false">https://www.sigarch.org/?p=108786</guid>
		<description><![CDATA[<div><img width="300" xheight="188" src="https://www.sigarch.org/wp-content/uploads/2026/07/mural-twilight-raleigh-convention-center-e1783651169742-300x188.jpg" class="attachment-medium size-medium wp-post-image" alt="" style="max-width:100% !important;height:auto !important;margin-bottom:15px;margin-left:15px;float:right;"  loading="lazy" /></div>The conference The 53rd International Symposium on Computer Architecture (ISCA) was held at the Raleigh Convention Center in Raleigh, North Carolina, from June 27 to July 1, 2026. Raleigh sits at one corner of the Research Triangle, anchored by North Carolina State University, the University of North Carolina at Chapel Hill, and Duke University. General [&#8230;]]]>
</description>
				<content:encoded><![CDATA[<div><img width="300" height="188" src="https://www.sigarch.org/wp-content/uploads/2026/07/mural-twilight-raleigh-convention-center-e1783651169742-300x188.jpg" class="attachment-medium size-medium wp-post-image" alt="" style="margin-bottom:15px;margin-left:15px;float:right;" decoding="async" loading="lazy" /></div><h3>The conference</h3>
<p><span style="font-weight: 400;">The 53rd International Symposium on Computer Architecture (ISCA) was held at the Raleigh Convention Center in Raleigh, North Carolina, from June 27 to July 1, 2026. Raleigh sits at one corner of the Research Triangle, anchored by North Carolina State University, the University of North Carolina at Chapel Hill, and Duke University. General Chairs Huiyang Zhou and James Tuck, both of NC State, led the organizing effort.</span></p>
<p><span style="font-weight: 400;">The most notable structural change this year was that ISCA offered remote attendance, making it a hybrid conference. The organizers provided deeply discounted remote registration to broaden access for students and researchers who could not travel, broadcast the main and keynote sessions on Zoom, and made recordings available to registrants for offline viewing. This was ISCA&#8217;s first hybrid offering and an experiment intended to lay groundwork for remote attendance at future architecture conferences.</span></p>
<h3></h3>
<h3>Workshops and tutorials</h3>
<p><span style="font-weight: 400;">Preceding the main symposium, ISCA 2026 opened with two full days of workshops and tutorials on Saturday, June 27 and Sunday, June 28, organized by Workshops and Tutorials Co-Chairs Lisa Wu Wills (Duke) and Brandon Reagen (NYU). The program totaled 15 workshops and 16 tutorials, spanning the full breadth of the field, from open-source infrastructure and DRAM to quantum computing, encrypted AI, and agentic design.</span></p>
<p><span style="font-weight: 400;">Saturday&#8217;s tutorials leaned on open-source and simulation infrastructure, such as</span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://astra-sim.github.io/tutorials/isca-2026"> <span style="font-weight: 400;">ASTRA-sim</span></a><span style="font-weight: 400;">, the</span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://events.safari.ethz.ch/isca26-ramulator-drambender/"> <span style="font-weight: 400;">Ramulator and DRAM Bender</span></a><span style="font-weight: 400;"> memory tools, and</span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://fava.stanford.edu/"> <span style="font-weight: 400;">FAVA</span></a><span style="font-weight: 400;"> on formal hardware verification, while the workshops included</span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://www.gem5.org/events/isca-2026"> <span style="font-weight: 400;">gem5</span></a><span style="font-weight: 400;">,</span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://cmu-caos.github.io/safeAI/2026/"> <span style="font-weight: 400;">SAFE AI</span></a><span style="font-weight: 400;"> on encrypted AI, and</span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://harvard-edge.github.io/isca-26-arch-2-workshop/"> <span style="font-weight: 400;">Architecture 2.0</span></a><span style="font-weight: 400;"> on agentic AI for computing-systems design, which marked the launch of the book </span><i><span style="font-weight: 400;">Architecture 2.0: Agentic Design Loops for Computing System Synthesis</span></i><span style="font-weight: 400;">. Sunday leaned into mentoring, open-source hardware, and quantum: the</span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://sites.google.com/view/uarchworkshop/home"> <span style="font-weight: 400;">uArch Mentoring Workshop</span></a><span style="font-weight: 400;"> and</span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://yarch2026.epfl.ch/"> <span style="font-weight: 400;">YArch&#8217;26</span></a><span style="font-weight: 400;"> for students;</span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://tutorial.xiangshan.cc/isca26/"> <span style="font-weight: 400;">XiangShan</span></a><span style="font-weight: 400;"> and the</span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://hpc.pnl.gov/SODA/tutorials/2026/ISCA2026.html"> <span style="font-weight: 400;">SODA Synthesizer</span></a><span style="font-weight: 400;"> on the open-source side; and</span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://scale.snu.ac.kr/isca2026-cheddar-tutorial/"> <span style="font-weight: 400;">FHE &amp; Cheddar</span></a><span style="font-weight: 400;"> and</span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://janusq.github.io/ISCA_2026_Tutorial/"> <span style="font-weight: 400;">Janus 4.0</span></a><span style="font-weight: 400;"> for quantum, alongside the 6th</span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://dramsec.ethz.ch/"> <span style="font-weight: 400;">DRAMSec</span></a><span style="font-weight: 400;">, a</span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://s4ai-cornelltech.github.io/ACT-ISCA/2026/"> <span style="font-weight: 400;">carbon-accounting</span></a><span style="font-weight: 400;"> tutorial, and the</span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://cbp-ng.bpchamp.com/"> <span style="font-weight: 400;">Championship in Branch Prediction</span></a><span style="font-weight: 400;">.</span></p>
<h3><img loading="lazy" decoding="async" class=" wp-image-108799 aligncenter" src="https://www.sigarch.org/wp-content/uploads/2026/07/IMG_2602-300x205.jpg" alt="" width="429" height="293" /></h3>
<p style="text-align: center;"><em><span style="font-weight: 400;">Panel on the Impact of AI on Higher Education &amp; Computer Architecture @uArch 2026</span></em></p>
<h3></h3>
<h3>The main program</h3>
<p><span style="font-weight: 400;">This was the largest ISCA program ever. Program Co-Chair Carole-Jean Wu (FAIR, Meta) and Kevin Skadron (University of Virginia) reported 850 regular-track submissions, a 49% increase over the previous year, of which 161 were accepted, for an 18.9% acceptance rate (down from 23% the year before). To accommodate the volume, ISCA ran a fourth parallel track for the first time. </span></p>
<p><span style="font-weight: 400;">The reviewing operation scaled to match. It involved 22 Area Chairs, 211 full PC members, and 192 lightweight PC members, by far the largest committee in the conference&#8217;s history. Reviewing ran in two rounds, with most papers reaching six reviews. Discussion followed the &#8220;Identify the Champion&#8221; model; 301 papers reached a clear online consensus, while the remaining 59 were resolved in a series of real-time Zoom PC meetings held over five days. In the end, 116 papers were accepted outright and another 45 were conditionally accepted with shepherding, all of which were eventually accepted.</span></p>
<h3></h3>
<h3>Keynotes</h3>
<p><span style="font-weight: 400;">ISCA 2026 featured three keynotes.</span></p>
<p><span style="font-weight: 400;">Debbie Marr (CEO and Co-Founder of AheadComputing) opened with &#8220;Computing at the Crossroads: Architecture, Economics, and the Next Era.&#8221; She reflected on the trajectories that shaped the field: Moore&#8217;s Law, Dennard scaling, increasing abstraction, and the long expansion of general-purpose computing. Many of those assumptions, she observed, are now being questioned simultaneously. She tied the technical inflection point to shifting economics, ecosystem dynamics, and leadership transitions, and suggested that the architecture community&#8217;s choices today will define the next era of computing.</span></p>
<p><span style="font-weight: 400;">The second keynote piloted a new &#8220;dialogue&#8221; format on quantum computing, pairing Fred Chong (University of Chicago; Chief Scientist for Quantum Software at Infleqtion) and Jay Gambetta (IBM Fellow and Director of Research) for &#8220;Architecting Hybrid Quantum-Classical Computing for Scale and Fault Tolerance.&#8221; Their shared theme: with fault-tolerant machines on the horizon and near-term machines increasingly integrated with classical HPC, computing will be heterogeneous and accelerator-based, and architects are needed to bridge theory and physical technology across applications, software, error correction, workflow management, and machine organization.</span></p>
<p><span style="font-weight: 400;">Babak Falsafi (EPFL) closed the keynote lineup with &#8220;Beyond the AI Energy Wall: Optimal Server Design and Operation&#8221;. He described how AI is pushing cloud infrastructure toward an energy wall, with compute demand growing faster than power, cooling, and datacenter capacity can be sustainably provisioned, and suggested that clearing it requires full-stack optimization rather than simply scaling accelerators or building larger facilities. He questioned the long-standing assumption that single-thread performance should dominate server design and operation.</span></p>
<h3><img loading="lazy" decoding="async" class="wp-image-108803 aligncenter" src="https://www.sigarch.org/wp-content/uploads/2026/07/IMG_2481-300x225.jpg" alt="" width="423" height="317" /></h3>
<p style="text-align: center;"><em>Keynote by Debbie Marr (Computing at the Crossroads)</em></p>
<h3></h3>
<h3>Awards</h3>
<p><span style="font-weight: 400;">A number of the community&#8217;s honors were presented during the conference.</span></p>
<p><span style="font-weight: 400;">The ACM/IEEE-CS Eckert-Mauchly Award went to Srinivas Devadas (MIT) for pioneering contributions to secure architectures with broad industrial and academic impact. The ACM SIGARCH Maurice Wilkes Award was presented to Tushar Krishna (Georgia Tech) for outstanding contributions to architectures and modeling tools for large-scale AI systems. The TCCA Young Architect Award went to Akshitha Sriraman (Carnegie Mellon University) for contributions to the design and management of efficient and sustainable cloud datacenters. The ACM SIGARCH/IEEE CS TCCA Outstanding Dissertation Award went to Olivia Hsu (Stanford University), with an honorable mention to Jovan Stojkovic (University of Illinois Urbana-Champaign). The SIGARCH Alan D. Berenbaum Distinguished Service Award was presented to Sarita Adve (University of Illinois Urbana-Champaign) for sustained and transformative contributions to ACM SIGARCH, the broader architecture community, and via CARES, the ACM SIG ecosystem. The ISCA Influential Paper Award recognized </span><i><span style="font-weight: 400;">&#8220;Adaptive Insertion Policies for High Performance Caching&#8221;</span></i><span style="font-weight: 400;"> (ISCA 2007) by Moinuddin K. Qureshi, Aamer Jaleel, Yale N. Patt, Simon C. Steely, and Joel Emer, for its commercial impact and for reinvigorating research on cache management with an elegant set-dueling framework that can be broadly applied to cache-policy selection.</span></p>
<p><span style="font-weight: 400;">Two ISCA Best Paper Awards were selected from a field of five nominations: </span><i><span style="font-weight: 400;">&#8220;Cerberus: Cross-Layer ECC Co-Design for Robust and Efficient Memory Protection&#8221;</span></i><span style="font-weight: 400;"> — Junhwan Kim, Seunghyun Kim, Yesin Ryu, Saeid Gorgin, and Jungrae Kim. </span><i><span style="font-weight: 400;">&#8220;Patterns Behind Chaos: Forecasting Data Movement for Efficient Large-Scale MoE LLM Inference&#8221;</span></i><span style="font-weight: 400;"> — Zhongkai Yu, Yue Guan, Zihao Yu, Chenyang Zhou, Zhengding Hu, Shuyi Pei, Yangwook Kang, Yufei Ding, and Po-An Tsai.</span></p>
<p><span style="font-weight: 400;">Two ISCA Distinguished Artifact Awards were also recognized: </span><i><span style="font-weight: 400;">&#8220;Transpiler-Architecture Co-Design to Curb Clifford Costs in Fault-Tolerant Quantum Computing&#8221;</span></i><span style="font-weight: 400;"> — Meng Wang, Chenxu Liu, Samuel Stein, Yufei Ding, Poulami Das, Prashant Nair, and Ang Li. </span><i><span style="font-weight: 400;">&#8220;Towards Practical Interrupt Side-Channel Attacks on macOS for Apple Silicon&#8221;</span></i><span style="font-weight: 400;"> — Xin Zhang, Chang Liu, Jiajun Zou, Yi Yang, Qingni Shen, Zhi Zhang, and Trevor E. Carlson.</span></p>
<h3><img loading="lazy" decoding="async" class="alignnone wp-image-108805 aligncenter" src="https://www.sigarch.org/wp-content/uploads/2026/07/IMG_2603-300x225.jpg" alt="" width="420" height="315" /></h3>
<p style="text-align: center;"><em>ISCA Influential Paper Award recipients</em></p>
<h3></h3>
<h3>Industry track and artifact evaluation</h3>
<p><span style="font-weight: 400;">The Industry Track, chaired by Brad Beckmann (AMD), accepted 11 papers out of 27, reviewed by a 32-member committee drawn entirely from industry across a diverse set of startups and established companies. The accepted set ranged from silicon to software. Two additional papers were recommended for an IEEE Micro Special Issue on Commercial Products.</span></p>
<p><span style="font-weight: 400;">Artifact Evaluation, in its fourth year at ISCA, received 49 submissions. 42 papers earned all three badges (Available, Functional, and Reproduced), 3 earned Available and Functional, and 4 earned Available. The co-chairs Hyeran Jeon (UC Merced), Linghao Song (Yale), and Mark Zhao (University of Colorado Boulder) flagged a growing challenge: the increasing heterogeneity of hardware and software platforms, which reviewers do not always have access to.</span></p>
<h3><em><img loading="lazy" decoding="async" class="alignnone wp-image-108804 aligncenter" src="https://www.sigarch.org/wp-content/uploads/2026/07/IMG_2491-300x223.jpg" alt="" width="425" height="316" /></em></h3>
<p style="text-align: center;"><em>A snapshot of the Industry Track session</em></p>
<h3></h3>
<h3>Excursion</h3>
<p><span style="font-weight: 400;">ISCA&#8217;s excursion was an evening dinner and social at Raleigh&#8217;s historic</span> <span style="font-weight: 400;">City Market</span><span style="font-weight: 400;">. Built in 1914 and known for its cobblestone streets and early-twentieth-century lamplight, the district hosted a relaxed, open-air affair, with food stations of North Carolina–inspired dishes, beer, and wine spread across the historic Market Hall, The Grove, and the outdoor spaces between them. With attendees spilling across the market, it made for an excellent networking opportunity and a welcome chance to unwind and catch up with people after the intensity of the technical program.</span></p>
<p><em><img loading="lazy" decoding="async" class="alignnone wp-image-108806 aligncenter" src="https://www.sigarch.org/wp-content/uploads/2026/07/IMG_2604-300x213.jpg" alt="" width="426" height="303" /></em></p>
<p style="text-align: center;"><em>Excursion venue: City Market</em></p>
<p>&nbsp;</p>
<p><b>About the author</b><span style="font-weight: 400;">: <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~libingyao.github.io">Bingyao Li</a> is an Assistant Professor in the Computer Science and Engineering Department at the University of California, Riverside. Her research focuses on designing architecture and system features for next-generation GPU platforms and building high-performance LLM infrastructure and systems.</span></p>
<p class="disclaim"><strong>Disclaimer:</strong> <em>These posts are written by individual contributors to share their thoughts on the Computer Architecture Today blog for the benefit of the community. Any views or opinions represented in this blog are personal, belong solely to the blog author and do not represent those of ACM SIGARCH or its parent organization, ACM.</em></p><Img align="left" border="0" height="1" width="1" alt="" style="border:0;float:left;margin:0;padding:0;width:1px!important;height:1px!important;" hspace="0" src="https://feeds.feedblitz.com/~/i/960054881/0/sigarch-cat">
]]>
</content:encoded>
			<wfw:commentRss>https://feeds.feedblitz.com/~/960054881/0/sigarch-cat/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
	<post-id xmlns="com-wordpress:feed-additions:1">108786</post-id></item>
<item>
<feedburner:origLink>https://www.sigarch.org/the-return-of-rigorous-full-system-timing-simulation/</feedburner:origLink>
		<title>The Return of Rigorous Full-System Timing Simulation</title>
		<link>https://feeds.feedblitz.com/~/957866780/0/sigarch-cat/</link>
		<comments>https://feeds.feedblitz.com/~/957866780/0/sigarch-cat/#respond</comments>
		<pubDate>Mon, 08 Jun 2026 15:00:17 +0000</pubDate>
		<dc:creator><![CDATA[Shanqing Lin, Mohammad Alian, Babak Falsafi]]></dc:creator>
		<category><![CDATA[ACM SIGARCH]]></category>
		<category><![CDATA[Simulation]]></category>
		<guid isPermaLink="false">https://www.sigarch.org/?p=105151</guid>
		<description><![CDATA[<div><img width="300" xheight="188" src="https://www.sigarch.org/wp-content/uploads/2026/05/ChatGPT-Image-May-30-2026-at-07_00_29-PM-300x188.png" class="attachment-medium size-medium wp-post-image" alt="" style="max-width:100% !important;height:auto !important;margin-bottom:15px;margin-left:15px;float:right;"  loading="lazy" /></div>The Return of Rigorous Full-System Timing Simulation Accurate timing simulation remains one of the most important tools in computer architecture, but modern systems have made cycle-level simulation increasingly impractical. Today’s platforms combine many-core CPUs, deep memory hierarchies, accelerators, complex I/O, and large software stacks, making detailed simulation extremely slow—often requiring months to simulate seconds of [&#8230;]]]>
</description>
				<content:encoded><![CDATA[<div><img width="300" height="188" src="https://www.sigarch.org/wp-content/uploads/2026/05/ChatGPT-Image-May-30-2026-at-07_00_29-PM-300x188.png" class="attachment-medium size-medium wp-post-image" alt="" style="margin-bottom:15px;margin-left:15px;float:right;" decoding="async" loading="lazy" /></div><h1><span style="font-weight: 400;">The Return of Rigorous Full-System Timing Simulation</span></h1>
<p><span style="font-weight: 400;">Accurate timing simulation remains one of the most important tools in computer architecture, but modern systems have made cycle-level simulation increasingly impractical. Today’s platforms combine many-core CPUs, deep memory hierarchies, accelerators, complex I/O, and large software stacks, making detailed simulation extremely slow—often requiring months to simulate seconds of execution. This “timing simulation wall” has pushed researchers toward approximations such as application-only simulation, fixed instruction windows, or instruction windows representing only the workload. While these reduce runtime, they often sacrifice rigorous end-to-end measurement of real microarchitectural behavior.</span></p>
<p><span style="font-weight: 400;">This blog argues for a return to rigorous full-system timing simulation—not by simulating everything in detail at all times, but by measuring the right execution intervals, using the right performance metrics, and applying statistically sound methods to make accurate simulation practical again.</span></p>
<h2><span style="font-weight: 400;">Why Full-System Simulation?</span></h2>
<p><span style="font-weight: 400;">Full-system simulation emulates an entire computer system: CPU, memory, devices, operating system, and applications. Unlike user-level simulation, it captures interactions across the full software and hardware stack. Full-system simulation matters because many critical behaviors emerge from OS activity, interrupts, I/O, memory management, synchronization, and device interactions—not from application code alone. Ignoring these layers can misrepresent real system bottlenecks and performance.</span></p>
<p><span style="font-weight: 400;">Full-system simulation dates back to the 1990s with systems like <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://en.wikipedia.org/wiki/SimOS">SimOS</a>, later influencing platforms such as <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://en.wikipedia.org/wiki/Simics">Simics</a> (now <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://www.intel.com/content/www/us/en/developer/articles/tool/simics-simulator.html">Intel Simics Simulator</a>), M5 (integrated into <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://dl.acm.org/doi/10.1145/2024716.2024718">gem5</a>) and <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://www.qemu.org">QEMU</a> (used in <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://ieeexplore.ieee.org/document/5982026">MARSS</a> and <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://qflex.epfl.ch">QFlex</a>).</span></p>
<p><span style="font-weight: 400;">Today, full-system simulation is becoming essential again for four reasons:</span></p>
<ol>
<li style="font-weight: 400;"><span style="font-weight: 400;">Modern workloads are service-oriented and multi-tenant, relying on microservices, RPCs, storage stacks, and OS-mediated interactions.</span></li>
<li style="font-weight: 400;"><span style="font-weight: 400;">Many server and mobile workloads spend significant time in the OS, making kernel behavior central to performance analysis.</span></li>
<li style="font-weight: 400;"><span style="font-weight: 400;">Heterogeneous systems increasingly combine CPUs with GPUs, accelerators, and smart NICs, with the CPU and OS orchestrating coordination, memory, and synchronization.</span></li>
<li style="font-weight: 400;"><span style="font-weight: 400;">Agentic AI workloads depend heavily on tool invocation, scheduling, APIs, databases, and system integration, making CPU and OS behavior critical to end-to-end performance.</span></li>
</ol>
<p><span style="font-weight: 400;">As a result, full-system simulation is no longer just a legacy methodology—it is increasingly necessary because the entire system stack has become the target of architectural innovation.</span></p>
<h1><span style="font-weight: 400;">The Timing Simulation Wall</span></h1>
<p><span style="font-weight: 400;">Simulators span a broad spectrum of abstraction, functionality, and performance. At the fastest end are execution-driven full-system simulators that use JIT translation to dynamically map target ISA instructions into the host ISA at runtime. Since early systems such as SimOS, these simulators have typically operated within roughly an order of magnitude of native hardware speed.</span></p>
<p><span style="font-weight: 400;">Modern ISA emulators such as QEMU can additionally generate detailed execution traces for functional simulation, enabling analysis of cache and TLB miss rates, branch predictor behavior, and prefetcher accuracy. This tracing introduces another order-of-magnitude slowdown relative to native execution.</span></p>
<p><span style="font-weight: 400;">Timing simulators go further by modeling cycle-level interactions among microarchitectural components in the CPU, accelerator, memory and I/O devices resulting in substantially lower simulation throughput. The table below compares simulation speeds for a single ARM Neoverse N1 target core with its cache hierarchy running server workloads on an AMD Zen 3 host.  The first row presents QEMU’s raw ISA emulation speed. The second row shows the slowdown due to instrumentation for user-level functional simulation. The third row demonstrates the impact on speed when functionally simulating the microarchitectural components, including the cache hierarchy and TLBs, front-end tables, and data prefetcher, for all user-level instructions. The fourth row shows the impact of functional simulation of all instructions, including the OS. Finally, the fifth row shows the timing simulation speed.</span><span style="font-weight: 400;"><img loading="lazy" decoding="async" class="wp-image-105366 aligncenter" src="https://www.sigarch.org/wp-content/uploads/2026/05/Screenshot-2026-05-26-at-2.53.13-PM-scaled.png" alt="" width="539" height="245" /></span></p>
<p><span style="font-weight: 400;">Modern workloads are not steady streams of similar instructions. Their performance fluctuates over time due to network activity, resource contention, background OS activity, synchronization effects, software hiccups, DVFS throttling, UI and graphics activity, and other asynchronous events. </span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://ieeexplore.ieee.org/abstract/document/1183520"><span style="font-weight: 400;">Alameldeen et al.</span></a><span style="font-weight: 400;"> presented a statistically rigorous methodology to determine the minimum measurement window needed to capture workload performance variability within a specified error bound and confidence level. </span></p>
<p><span style="font-weight: 400;">Unlike conventional database workloads (e.g., TPC benchmarks) which have prescribed measurement windows, typical benchmarks and workloads used in research do not. Applying Alameldeen’s methodology, we find that capturing performance variability for a single ARM Neoverse N1 core and its cache hierarchy requires five to 120 seconds of target execution time across server workloads from CloudSuite, DCPerf, and DeathStarBench. Simulating even a few seconds of a single core with today’s fastest cycle-accurate simulator, gem5, at 250 KIPS requires months of simulation time.</span></p>
<h2><span style="font-weight: 400;">What Should We Measure?</span></h2>
<p><span style="font-weight: 400;">The second question is which performance metric to use. Timing simulators count cycles, so architects often report IPC, or instructions per cycle. IPC is reasonable for single-core workloads when most executed instructions correspond to program progress.</span></p>
<p><span style="font-weight: 400;">For multicore workloads, however, </span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://dl.acm.org/doi/10.1109/MM.2006.73"><span style="font-weight: 400;">IPC can be misleading</span></a><span style="font-weight: 400;">. Threads may spin, poll, block, wait on locks, synchronize, or execute OS code that does not advance useful work. A system can therefore sustain high IPC while making little forward progress; in effect, total IPC can reward busy waiting. </span></p>
<p><span style="font-weight: 400;">This is why </span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://ieeexplore.ieee.org/document/1677500"><span style="font-weight: 400;">user-level IPC</span></a><span style="font-weight: 400;">, or U-IPC, is often a better proxy. U-IPC counts user-level instructions over time, assuming that user instructions per request remain roughly stable and that most spinning occurs in the OS. Under that assumption, U-IPC tracks useful throughput more directly than total IPC.</span></p>
<p><span style="font-weight: 400;">But U-IPC must be validated for each workload. If spinning occurs in user space, as in systems with user-level network stacks, raw U-IPC still counts non-productive work and must be corrected to exclude spinning. The broader requirement is therefore metric validation: a rigorous simulation methodology must show that the chosen metric—IPC, U-IPC, throughput, or latency—actually captures forward progress for the workload under study.</span></p>
<h1><span style="font-weight: 400;">How Should We Measure?</span></h1>
<p><span style="font-weight: 400;">Due to the timing simulation wall, researchers often use abbreviated measurements. The most common technique is to measure a single unit of 100 million to one billion instructions. Unfortunately, depending on where in the execution the fixed measurement is taken from, this technique may lead to inconclusive results or worse, incorrect conclusions. </span></p>
<p><span style="font-weight: 400;">Instead, designers often use sampling to capture variability in performance estimates. Phase-based sampling, such as </span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://dl.acm.org/doi/10.1145/885651.781076"><span style="font-weight: 400;">SimPoint</span></a><span style="font-weight: 400;">, is a popular technique that uses clustering of basic-block vectors (BBVs) to select representative application “phases.” Such sampling properly captures the representing repetitive instruction streams that account for most of the execution. </span></p>
<p><span style="font-weight: 400;">While simple and practical, phase-based sampling may ignore OS effects, interrupts and I/O interactions, communication among cores, and software hiccups. Moreover, </span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://users.ece.cmu.edu/~jhoe/distribution/2010/wunderlich.pdf"><span style="font-weight: 400;">Wunderlich</span></a><span style="font-weight: 400;"> argues in his thesis that phase-based sampling: (1) misses the microarchitectural footprint of less common instruction streams and their impact on performance, and (2) forgoes any error bounds with confidence in estimates. </span></p>
<p><span style="font-weight: 400;">A rigorous sampling technique is statistical sampling, such as </span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://ieeexplore.ieee.org/document/1206991"><span style="font-weight: 400;">SMARTS</span></a><span style="font-weight: 400;">, taking a large sample (e.g., hundreds) of small (e.g., 200k cycles), uniformly distanced measurement units that is representative of execution, not phases in the workload. This technique enables bounding the error in estimates and delivers quantifiable confidence. It also opens an entire plethora of statistical sampling tools to trade off confidence in estimates for measurement in speed and quantify sample divergence to detect bias in estimates.</span></p>
<p><span style="font-weight: 400;">The figure below compares error magnitude in performance estimates among three abbreviated measurement techniques from full-timing simulation runs of tens of target seconds on a two-core socket with 2.0 GHz ARM Neoverse N1 cores running single-tier, multi-tier and consolidated server workloads (CloudSuite, DCPerf and DeathStarBench). The figure compares the error against the full-timing baselines for: (1) single units of one billion instructions per core starting from three equally distanced positions in the minimum measurement window (i.e., beginning, 1/3 and 2/3 into the population), (2) units of 100 target microseconds (i.e., 200k cycles for a 2.0 GHz clock) including basic-block vectors (BBV) derived from K-means clustering, and (3) a uniform sample (of hundreds) of 100 target microseconds drawn with an error bound of 5% with 95% confidence with statistical sampling. </span></p>
<p><img loading="lazy" decoding="async" class="wp-image-105321 aligncenter" src="https://www.sigarch.org/wp-content/uploads/2026/05/Screenshot-2026-05-25-at-9.00.17-PM-scaled.png" alt="" width="521" height="264" /></p>
<p><span style="font-weight: 400;">Both one-billion instruction units and BBV result in high error estimates with the former not being representative of execution and the latter representing only frequently executed instructions. In contrast, statistical sampling results in a desired error bound with confidence because it represents not just frequently executed instructions but also instructions that have a high impact on performance due to their microarchitectural footprint.</span></p>
<h2><span style="font-weight: 400;">A SOTA Sampling Framework</span></h2>
<p><span style="font-weight: 400;">The figure below presents a state-of-the-art sampling framework using </span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://qflex.epfl.ch/"><span style="font-weight: 400;">QFlex 3.0</span></a><span style="font-weight: 400;"> (derived from </span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://dl.acm.org/doi/abs/10.1109/MM.2006.79"><span style="font-weight: 400;">SimFlex</span></a><span style="font-weight: 400;">) for full-system timing simulation of ARM ISA. For each workload, the software stack together with the OS is first loaded and warmed on a real platform, then tested to identify the minimum window&#8212;such as five to 120 target machine’s seconds&#8212;called a “population”, using </span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://ieeexplore.ieee.org/abstract/document/1183520"><span style="font-weight: 400;">Alameldeen et al.</span></a><span style="font-weight: 400;">’s technique. The workload is then loaded again, this time with QEMU and run through a functional simulator running on average at 6 MIPS for the entire duration of population. </span></p>
<p><img loading="lazy" decoding="async" class="wp-image-105319 aligncenter" src="https://www.sigarch.org/wp-content/uploads/2026/05/Screenshot-2026-05-25-at-3.15.58-PM.png" alt="" width="536" height="298" /></p>
<p><span style="font-weight: 400;">The functional simulator simulates all microarchitectural components with long-term state (e.g., cache hierarchy and TLBs, branch tables, data prefetcher) and periodically dumps checkpoints with architectural and microarchitectural state into a checkpoint library. Because the functional simulator is not cycle-accurate, it requires an approximation for time. The most common approximation is </span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://dl.acm.org/doi/10.1145/2063384.2063454"><span style="font-weight: 400;">IPC=1 </span></a><span style="font-weight: 400;">or IPC derived from </span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://ieeexplore.ieee.org/abstract/document/6522340"><span style="font-weight: 400;">neighboring units</span></a><span style="font-weight: 400;">.</span></p>
<p><span style="font-weight: 400;">The checkpoints in the library are then run using a timing simulator for 100 us independently and embarrassingly parallel. Each checkpoint is first run for a bounded window of time (e.g., 200 us) to make sure microarchitectural components with short-term state (e.g., buffers in the pipeline, cache hierarchy and NoC) are warm, followed by a measurement. The timing results for the sample are then aggregated to determine whether the sample (i.e., number of checkpoints) is large enough to bound the error for a desired level of confidence (e.g., 5% with 95% confidence). If not, the sampling framework creates a new checkpoint library with a shorter interval between checkpoints.</span></p>
<p><span style="color: #333333; font-size: 26px;">Challenges and Open Problems </span></p>
<p><span style="font-weight: 400;">Even with accurate measurement techniques, there are fundamental challenges with sampling (for both phase-based and statistical sampling).</span></p>
<ol>
<li style="font-weight: 400;"><b>Accurate state generation. </b><span style="font-weight: 400;">Timing-induced activity during functional simulation and its impact on the microarchitectural footprint may result in a significant bias because time is approximated. This challenge is more pronounced with variable performance among target threads in multi-tier and consolidated workloads where the speed bias may significantly impact the resulting shared microarchitectural footprint.</span></li>
<li style="font-weight: 400;"><b>The functional simulation wall.</b><span style="font-weight: 400;"> Sampling minimizes the required measurement using timing simulators but shifts the bottleneck to the functional simulator (which at 6 MIPS is only 24x faster than a 250 KIPs timing simulator). Parallelizing functional simulation may be a promising approach to enable scalability with multicore hosts. Parallel simulation is fundamentally limited by the granularity at which target threads communicate.</span></li>
<li style="font-weight: 400;"><b>Support for checkpointing.</b><span style="font-weight: 400;"> Generating and restoring an entire checkpoint for every measurement is impractical in both storage capacity and runtime overhead. Practical sampling therefore requires incremental checkpoint storage and restoration.</span></li>
<li style="font-weight: 400;"><b>Sampling non-average metrics.</b><span style="font-weight: 400;"> Statistical sampling works well for average-like metrics such as IPC or U-IPC, but it is harder to apply to extreme or rare-event metrics such as maximum </span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://ieeexplore.ieee.org/abstract/document/6522340"><span style="font-weight: 400;">temperature</span></a><span style="font-weight: 400;">, </span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://users.ece.cmu.edu/~jhoe/distribution/2010/wunderlich.pdf"><span style="font-weight: 400;">worst-case power</span></a><span style="font-weight: 400;">, or rare latency spikes.</span></li>
<li style="font-weight: 400;"><b>Capturing service-level metrics.</b><span style="font-weight: 400;"> Metrics such as request latency or p99.9 latency are much coarser-grained than sampling units needed for IPC or U-IPC. Capturing service-level metrics and tail latency may require an order of magnitude larger population and sampling units which poses a challenge for both functional and timing simulation.</span></li>
<li style="font-weight: 400;"><b>Multi-node full-system simulation.</b><span style="font-weight: 400;"> Many modern workloads are distributed across multiple machines. Single-node simulation is often insufficient for datacenter-scale behavior, but despite </span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://ieeexplore.ieee.org/document/7975287"><span style="font-weight: 400;">progress</span></a><span style="font-weight: 400;">, rigorous  multi-node full-system timing simulation remains an open challenge.</span></li>
<li style="font-weight: 400;"><b>Interoperability across simulators.</b><span style="font-weight: 400;"> A practical ecosystem should allow one tool to generate a checkpoint library and another to perform timing simulation. This interoperability requires an interface definition language allowing interoperable architectural and microarchitectural state among simulators.</span></li>
</ol>
<h2><span style="font-weight: 400;">About the Authors</span></h2>
<p><span style="font-weight: 400;"><strong>Shanqing Lin</strong> is a final-year PhD student at the School of Computer and Communication Sciences at EPFL and the principal developer of QFlex v3.0.</span></p>
<p><span style="font-weight: 400;"><strong>Mohammad Alian</strong> is an Assistant Professor in the Electrical and Computer Engineering Department at Cornell University.</span></p>
<p><span style="font-weight: 400;"><strong>Babak Falsafi</strong> is a Professor in the School of Computer and Communication Sciences at EPFL (epfl.ch) and the founding President of Swiss Datacenter Efficiency Association (sdea.ch).</span></p>
<p class="disclaim"><strong>Disclaimer:</strong> <em>These posts are written by individual contributors to share their thoughts on the Computer Architecture Today blog for the benefit of the community. Any views or opinions represented in this blog are personal, belong solely to the blog author and do not represent those of ACM SIGARCH or its parent organization, ACM.</em></p><Img align="left" border="0" height="1" width="1" alt="" style="border:0;float:left;margin:0;padding:0;width:1px!important;height:1px!important;" hspace="0" src="https://feeds.feedblitz.com/~/i/957866780/0/sigarch-cat">
]]>
</content:encoded>
			<wfw:commentRss>https://feeds.feedblitz.com/~/957866780/0/sigarch-cat/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
	<post-id xmlns="com-wordpress:feed-additions:1">105151</post-id></item>
<item>
<feedburner:origLink>https://www.sigarch.org/agentic-security-lessons-from-computer-architecture/</feedburner:origLink>
		<title>Agentic Security: Lessons from Computer Architecture</title>
		<link>https://feeds.feedblitz.com/~/957647303/0/sigarch-cat/</link>
		<comments>https://feeds.feedblitz.com/~/957647303/0/sigarch-cat/#respond</comments>
		<pubDate>Tue, 02 Jun 2026 14:05:33 +0000</pubDate>
		<dc:creator><![CDATA[Simha Sethumadhavan]]></dc:creator>
		<category><![CDATA[ACM SIGARCH]]></category>
		<category><![CDATA[AI Agents]]></category>
		<category><![CDATA[Security]]></category>
		<category><![CDATA[Spectre]]></category>
		<guid isPermaLink="false">https://www.sigarch.org/?p=105511</guid>
		<description><![CDATA[<div><img width="300" xheight="188" src="https://www.sigarch.org/wp-content/uploads/2026/05/ChatGPT-Image-May-29-2026-06_39_08-AM-300x188.png" class="attachment-medium size-medium wp-post-image" alt="" style="max-width:100% !important;height:auto !important;margin-bottom:15px;margin-left:15px;float:right;"  loading="lazy" /></div>When an agent makes an incorrect guess, the obvious mistakes like bad files or stale outputs are straightforward to see. However, there are less visible leaks that pose significant risks, such as timing patterns or cached context. The context and data exchanged between tools, services, and third-party systems can also be problematic. This situation becomes particularly concerning when AI agents take action before fully understanding the task at hand. This leads to an important question: Who holds the responsibility for addressing the residue left behind by agentic mistakes?
]]>
</description>
				<content:encoded><![CDATA[<div><img width="300" height="188" src="https://www.sigarch.org/wp-content/uploads/2026/05/ChatGPT-Image-May-29-2026-06_39_08-AM-300x188.png" class="attachment-medium size-medium wp-post-image" alt="" style="margin-bottom:15px;margin-left:15px;float:right;" decoding="async" loading="lazy" /></div><p id="ember54" class="ember-view reader-text-block__paragraph">What does speculative execution in a processor &#8212; and the predictor that drives it, such as a branch predictor &#8212; have to do with AI agents? They may <em>seem </em>very different, yet, at a high level of abstraction there are similarities.</p>
<p id="ember55" class="ember-view reader-text-block__paragraph">Both speculate: A processor predicts which way a branch will go and begins executing instructions along the predicted path before the branch has resolved. An AI agent infers a user’s intent, reads/writes files, executes programs, makes network calls etc., before it knows whether its interpretation of the user’s intent is right.</p>
<p id="ember56" class="ember-view reader-text-block__paragraph">Because prediction can fail, both systems require roll back mechanisms. In a processor, once a misprediction is detected, the wrong path work disappears (from the programmer&#8217;s point of view). When the agent is told it is wrong, or figures that out itself, it may revise its plan and redo the work after rolling back to a good checkpoint.</p>
<p id="ember57" class="ember-view reader-text-block__paragraph">Both systems litter and leave residues: While a processor can recover from a misprediction without any programmer visible effects, under the covers, microarchitecturally, wrong path execution perturbs on chip structures like caches. The Spectre attack (2018) showed that this residue can be observed through covert channels. AI agents have a similar problem. When an agent, or its human user, notices a mistake and corrects it, the failed attempt can leave at least two types of residues: a) residue that is easily observable like bad outputs, stale files or processes, or b) harder to know/track/undo residue like timing and volume of network requests, model context summaries shipped to third party servers to name a few.</p>
<p id="ember58" class="ember-view reader-text-block__paragraph">Also both systems can be tricked and steered through adversarial inputs: In Spectre, the attacker influences the on chip predictor state by executing a pattern, then supplies an adversarial input that causes the victim to transiently execute along the trained path that it should not take architecturally. While that transient execution is later squashed the microarchitectural residue of the execution can still be measured. Malicious prompts can play a similar role in AI agents: they can steer the system toward actions that are later corrected or denied, but in the process may leave litter data that attackers can use.</p>
<p id="ember59" class="ember-view reader-text-block__paragraph">Given these similarities, we can ask two questions.</p>
<p id="ember60" class="ember-view reader-text-block__paragraph">1) Can AI agents completely eliminate easily observable &#8220;architectural&#8221; residues on mispredictions?</p>
<p id="ember61" class="ember-view reader-text-block__paragraph">2) What are the dangers/risks of hidden &#8220;microarchitectural&#8221; residue left behind by AI agents?</p>
<p id="ember62" class="ember-view reader-text-block__paragraph">Regarding architectural residue, processors can hide speculative wrong path work cleanly because the ISA defines what counts as visible committed state. Currently there is no equivalent for AI agents: the absence of an interface that can precisely define operations, state, life time of state, and triggers for misprediction recovery, makes these systems hard to reason about and a fertile ground for leakage.</p>
<p id="ember63" class="ember-view reader-text-block__paragraph">While observable residue is a serious problem it is also a solvable problem to some degree: if one is satisfied with imprecise, best effort work, one simple thing to do is to just prompt the agent to clean up after itself. A really smart agent, <em>in theory</em>, should be able to use mechanisms like transactions, two phase commit, distributed undo protocols, disposable containers and VMs, sandboxes, access controls, versioning and information flow tracking to minimize overt residues. However, if we wanted to do better than prompting we probably will need an ISA-like layer.</p>
<p id="ember64" class="ember-view reader-text-block__paragraph">The second, and harder, question is about what happens to hidden/microarchitectural residues. In general, clean up of this type of residue is hard because it is often left in places no one thinks to inspect, or in places users cannot practically inspect because the those parts are proprietary or distributed across organizational boundaries. Also with AI agents, the problem is broader in scope than in a processor because it spans a larger number of tech layers from model context to hardware, local and remote. Further, an agent’s speculation window may last seconds or minutes, compared with nanoseconds in a processor. That longer window creates more opportunity for residue to diffuse. It is highly unlikely that we can simply prompt the agent to clean up hidden/microarchitectural residue because, by definition, there isn&#8217;t an architectural interface to observe or control microarchitectural state/work.</p>
<p id="ember65" class="ember-view reader-text-block__paragraph">How likely are we to solve agentic littering? Who needs this problem solved? And, who should solve this problem?</p>
<p id="ember66" class="ember-view reader-text-block__paragraph">In addition to technical aspects, economics and incentives often determine whether solutions are adopted. Here too we can look at processor misprediction recovery and compare them to AI agents.</p>
<p id="ember67" class="ember-view reader-text-block__paragraph">Overt architectural and hidden microarchitectural residues have different economics and incentives at play.</p>
<p id="ember68" class="ember-view reader-text-block__paragraph">Overt residues are easier to price. If an agent leaves behind a directory full of junk, consumes too many resources, or corrupts a file, that failure is visible to users. Users will complain, and because there are complaints, product teams can justify spending resources to fix them.</p>
<p id="ember69" class="ember-view reader-text-block__paragraph">Hidden residues are harder. These residues may not produce an obvious effect like a crash. They may also require complex conditions to manifest. That makes it harder to attribute with accuracy and consequently easier to dismiss. It also makes it harder for users to demand fixes, because users often cannot see the thing they are supposed to complain about.</p>
<p id="ember70" class="ember-view reader-text-block__paragraph">Spectre, an issue due to adversarial steering and microarchitectural residue, was disclosed roughly eight years ago, and the broader class of this leakage has still not been completely fixed. This is not because principled technical solutions do not exist. It is because these solutions increase design complexity, impact performance, change the hardware and software interface in ways that is not easy to adopt, or require coordination across vendors and different layers of the computing stack all of which add recurring or non-recurring costs. Also, each layer can plausibly say that the residue cleanup should be handled by someone else. Vendors can also say that there are have not seen large scale attacks and that they do not have to protect against these attacks given the risk profile.</p>
<p id="ember71" class="ember-view reader-text-block__paragraph">The same pattern may emerge for AI agents and handling hidden/microarchitectural residues.</p>
<p id="ember72" class="ember-view reader-text-block__paragraph">Each AI agent boundary is also an economic boundary. Each layer can plausibly say that the residue is someone else’s problem. The model provider can say the deployment should isolate side effects. The deployment/orchestrator can say the runtime should enforce cleanup. The runtime can say the operating system should provide better isolation. The hardware vendor can say software should avoid sensitive colocation. The user experiences the combined risk of all these but usually has the least ability to inspect or repair it!</p>
<p id="ember73" class="ember-view reader-text-block__paragraph">The real answer is that every party involved in agentic execution should fix its own leaks and share responsibility for security and privacy. But each party also has reason to argue that the cost is too high, especially when the economic benefits are difficult to measure and the harms are difficult to attribute.</p>
<p id="ember74" class="ember-view reader-text-block__paragraph">So the likely outcome here is not hard to guess. Hidden microarchitectural residue handling is treated as an afterthought, and agents end up reflecting the incentives that shaped it, viz., agents get more capable, overt residue cleanup improves through ad hoc clean up attempts, and create a very long tail of hard to detect, microarchitectural residues that expands the attack surface.</p>
<p id="ember75" class="ember-view reader-text-block__paragraph">The best chance for security is while these systems are being designed and deployed. AI-agent platforms designed now should at least treat residue management as first-class design requirement. That, however, means finding ways to incentivize designers to care about hidden, microarchitectural residue before users are harmed. If we treat microrchitectural residue management as an optional, &#8220;nice-to-have&#8221;, &#8220;less-important-than-overt&#8221; security feature, we will spend the next decade patching a massive, distributed attack surface.</p>
<p><strong>About the Author:</strong> Simha Sethumadhavan is a Professor in the CS department at Columbia University. He would like to thank  Profs. <a id="ember77" class="ember-view" href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://www.linkedin.com/in/roxana-geambasu-93b58b1b4/">Roxana Geambasu</a>. Martha Kim and <a id="ember78" class="ember-view" href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://www.linkedin.com/in/takhandipu/">Tanvir Ahmed Khan </a>for thought provoking comments and feedback.</p>
<p class="disclaim"><strong>Disclaimer:</strong> <em>These posts are written by individual contributors to share their thoughts on the Computer Architecture Today blog for the benefit of the community. Any views or opinions represented in this blog are personal, belong solely to the blog author and do not represent those of ACM SIGARCH or its parent organization, ACM.</em></p><Img align="left" border="0" height="1" width="1" alt="" style="border:0;float:left;margin:0;padding:0;width:1px!important;height:1px!important;" hspace="0" src="https://feeds.feedblitz.com/~/i/957647303/0/sigarch-cat">
]]>
</content:encoded>
			<wfw:commentRss>https://feeds.feedblitz.com/~/957647303/0/sigarch-cat/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
	<post-id xmlns="com-wordpress:feed-additions:1">105511</post-id></item>
<item>
<feedburner:origLink>https://www.sigarch.org/architecture-systems-are-changing-the-architects-role-in-the-era-of-agentic-co-design/</feedburner:origLink>
		<title>Architecture &#038; Systems are Changing: The Architect&#8217;s Role in the Era of Agentic Co-Design</title>
		<link>https://feeds.feedblitz.com/~/956665028/0/sigarch-cat/</link>
		<comments>https://feeds.feedblitz.com/~/956665028/0/sigarch-cat/#respond</comments>
		<pubDate>Tue, 19 May 2026 14:00:32 +0000</pubDate>
		<dc:creator><![CDATA[Dimitrios Skarlatos]]></dc:creator>
		<category><![CDATA[ACM SIGARCH]]></category>
		<category><![CDATA[AI Agents]]></category>
		<category><![CDATA[Hardware-Software Co-design]]></category>
		<guid isPermaLink="false">https://www.sigarch.org/?p=104529</guid>
		<description><![CDATA[<div><img width="300" xheight="200" src="https://www.sigarch.org/wp-content/uploads/2026/05/feature-300x200.png" class="attachment-medium size-medium wp-post-image" alt="" style="max-width:100% !important;height:auto !important;margin-bottom:15px;margin-left:15px;float:right;"  loading="lazy" /></div>Architecture &#38; Systems are Changing: The Architect&#8217;s Role in the Era of Agentic Co-Design The AI datacenter stack is built on hardware-software contracts and abstractions that were never designed for the workloads datacenters now serve. Memory systems strain under terabyte-scale capacity. Heterogeneous accelerators have been pressed into deployment. With datacenters projected to consume over 1,000 [&#8230;]]]>
</description>
				<content:encoded><![CDATA[<div><img width="300" height="200" src="https://www.sigarch.org/wp-content/uploads/2026/05/feature-300x200.png" class="attachment-medium size-medium wp-post-image" alt="" style="margin-bottom:15px;margin-left:15px;float:right;" decoding="async" loading="lazy" /></div><h1><b>Architecture &amp; Systems are Changing: The Architect&#8217;s Role in the Era of Agentic Co-Design</b></h1>
<p><span style="font-weight: 400;">The AI datacenter stack is built on hardware-software contracts and abstractions that were never designed for the workloads datacenters now serve. Memory systems strain under terabyte-scale capacity. Heterogeneous accelerators have been pressed into deployment. With datacenters projected to consume over 1,000 TWh annually, surpassing Japan (the world&#8217;s fourth-largest economy), renegotiating the hardware-software contract is no longer optional.</span></p>
<p><span style="font-weight: 400;">AI was enabled by decades of hardware and software efficiency gains. The next leap requires two orders of magnitude more, on a stack whose workloads, infrastructure, and economics bear little resemblance to the one the contract was written for.</span></p>
<p><span style="font-weight: 400;">That is not a problem any single layer of the stack can solve. It is a co-design problem, and it is unfolding while the design process itself is changing across systems and architecture.</span></p>
<h2><b>The contract so far</b></h2>
<p><span style="font-weight: 400;">Computer architecture has long been guided by a quiet contract with three commitments: </span><b>abstractions, interfaces, and transparency</b><span style="font-weight: 400;">. Layers that hide hardware complexity from programmers; interfaces like the x86 ISA that let decades-old binaries still run on Linux today; and microarchitectural state largely hidden behind a model programmers can keep in their heads. Together, these commitments deliver the property programmers care about most: </span><b>programmability</b><span style="font-weight: 400;">.</span></p>
<p><span style="font-weight: 400;">This contract is not arbitrary: it is what lets billions of lines of legacy software keep running while architects rebuild underneath. But the contract was negotiated for a world where humans wrote all of the code and humans designed all of the hardware. Both halves of that world are changing at the same time, and the architect&#8217;s job is evolving with them.</span></p>
<h2><b>Plenty of room at the Top</b></h2>
<p><span style="font-weight: 400;">In 2020, Leiserson, Thompson, Emer, Kuszmaul, Lampson, Sanchez, and Schardl argued in </span><i><span style="font-weight: 400;">Science</span></i><span style="font-weight: 400;"> that post-Moore performance gains would have to come from the</span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://doi.org/10.1126/science.aam9744"> <span style="font-weight: 400;">&#8220;Top&#8221; of the computing stack</span></a><span style="font-weight: 400;">: software, algorithms, and hardware architecture, rather than from the &#8220;Bottom&#8221; of semiconductor physics. They were right, and the half-decade since has only sharpened the point.</span></p>
<p><span style="font-weight: 400;">The harder claim in that paper is the one we want to dwell on. The Top has plenty of room, but the gains are </span><i><span style="font-weight: 400;">&#8220;opportunistic, uneven, and sporadic,&#8221;</span></i><span style="font-weight: 400;"> subject to diminishing returns. The Top has historically been mined by hand, one paper and one design cycle at a time. What is changing now is the rate at which it is </span><i><span style="font-weight: 400;">mineable</span></i><span style="font-weight: 400;">. The two directions we describe next change that rate. Same Top, mined faster, mined more systematically, and mined by tools the field did not have until recently.</span></p>
<h2><b>Two directions are reshaping the design loop</b></h2>
<p><span style="font-weight: 400;">Two complementary directions are converging on how we build system software and hardware: </span><b>embedding learning inside low-level mechanisms</b><span style="font-weight: 400;">, and </span><b>using AI agents to explore the architectural design space itself</b><span style="font-weight: 400;">.</span></p>
<p><span style="font-weight: 400;">The first direction has a deep history. Perceptron-based</span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://ieeexplore.ieee.org/document/903263"> <span style="font-weight: 400;">branch predictors</span></a><span style="font-weight: 400;"> put a lightweight learning model on the critical path more than two decades ago, and the catalog has steadily grown since. On the cache-hierarchy side,</span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://ieeexplore.ieee.org/document/9773195"> <span style="font-weight: 400;">Mockingjay</span></a><span style="font-weight: 400;"> uses a trained reuse-distance predictor to imitate Belady&#8217;s optimal replacement policy. On the prefetching side,</span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://arxiv.org/abs/1803.02329"> <span style="font-weight: 400;">Hashemi et al.</span></a><span style="font-weight: 400;"> framed memory access patterns as an LSTM prediction task,</span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://dl.acm.org/doi/10.1145/3466752.3480114"> <span style="font-weight: 400;">Pythia</span></a><span style="font-weight: 400;"> recast the entire prefetcher as an online reinforcement-learning agent, and</span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://dl.acm.org/doi/10.1145/3613424.3623780"> <span style="font-weight: 400;">Micro-Armed Bandit</span></a><span style="font-weight: 400;"> showed that lightweight bandit-based RL can match more complex agents at a fraction of the storage cost. Outside the cache hierarchy, reinforcement learning has been applied to</span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://www.nature.com/articles/s41586-021-03544-w"> <span style="font-weight: 400;">chip floorplanning</span></a><span style="font-weight: 400;">, learning-based</span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://dl.acm.org/doi/abs/10.1145/3373376.3378525"> <span style="font-weight: 400;">memory allocation</span></a><span style="font-weight: 400;"> replaced hand-tuned allocator heuristics with predictors trained on real telemetry, and</span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://dl.acm.org/doi/10.1145/3297858.3304004"> <span style="font-weight: 400;">Seer</span></a><span style="font-weight: 400;"> applied deep learning to predict QoS violations in cloud microservices before they materialize. Most recently, our work on learned virtual memory (</span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://dl.acm.org/doi/10.1145/3725843.3756093"><span style="font-weight: 400;">LVM</span></a><span style="font-weight: 400;">) eliminated address-translation overhead with a learned index that fits in two cycles of integer arithmetic. The principle generalizes: fixed designs are being replaced with principled, hardware-realizable models that adapt to workload shifts in ways hand-tuned heuristics cannot.</span></p>
<p><span style="font-weight: 400;">The second direction is newer, and arguably more disruptive.</span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://arxiv.org/abs/2506.13131"> <span style="font-weight: 400;">AlphaEvolve</span></a><span style="font-weight: 400;"> demonstrated that LLMs paired with evolutionary search can discover algorithms across domains, from mathematical constructions to data-center scheduling.</span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://arxiv.org/abs/2510.06189"> <span style="font-weight: 400;">ADRS</span></a><span style="font-weight: 400;"> extended the idea to broader systems research, and</span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://arxiv.org/abs/2602.22425"> <span style="font-weight: 400;">recent work from Google</span></a><span style="font-weight: 400;"> has applied the same approach to cache replacement. The same paradigm has reached the software side of the machine: agentic systems that generate and tune CUDA and Triton kernels, including NVIDIA’s </span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://arxiv.org/pdf/2603.24517"><span style="font-weight: 400;">AVO</span></a><span style="font-weight: 400;"> and Meta&#8217;s</span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://arxiv.org/abs/2512.23236"> <span style="font-weight: 400;">KernelEvolve</span></a><span style="font-weight: 400;">, are now in use across heterogeneous accelerators. Sankaralingam captured the bigger picture in</span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://arxiv.org/abs/2604.03312"> <i><span style="font-weight: 400;">Computer Architecture&#8217;s AlphaZero Moment</span></i></a><span style="font-weight: 400;">, arguing that the field is approaching a regime where architectural </span><i><span style="font-weight: 400;">discovery itself</span></i><span style="font-weight: 400;"> becomes a search problem, beyond per-mechanism tuning. In our own recent work,</span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://arxiv.org/abs/2604.25083"> <i><span style="font-weight: 400;">Agentic Architect</span></i></a><span style="font-weight: 400;">, we coupled LLM-driven code evolution with cycle-accurate simulation to explore microarchitectural design spaces, and found that the loop matches or exceeds state-of-the-art designs on cache replacement, prefetching, and branch prediction. </span></p>
<p><span style="font-weight: 400;">These two directions are not alternatives. They differ in what they decide and when. The first decides </span><b>how a fixed mechanism behaves at runtime</b><span style="font-weight: 400;">: a branch predictor that learns its own weights, a cache policy that adapts to the workload. The second decides </span><b>what the mechanism looks like in the first place</b><span style="font-weight: 400;">: the predictor, the policy, the prefetcher itself, evolved before deployment. Both move judgment that used to live in tight loops written by experts into search problems that can be scored and re-evaluated.</span></p>
<div id="attachment_104533" style="width: 2570px" class="wp-caption aligncenter"><img loading="lazy" decoding="async" aria-describedby="caption-attachment-104533" class="wp-image-104533 size-full" src="https://www.sigarch.org/wp-content/uploads/2026/05/framework-scaled.png" alt="" width="2560" height="784" srcset="https://www.sigarch.org/wp-content/uploads/2026/05/framework-scaled.png 2560w, https://www.sigarch.org/wp-content/uploads/2026/05/framework-1280x392.png 1280w, https://www.sigarch.org/wp-content/uploads/2026/05/framework-980x300.png 980w, https://www.sigarch.org/wp-content/uploads/2026/05/framework-480x147.png 480w" sizes="(min-width: 0px) and (max-width: 480px) 480px, (min-width: 481px) and (max-width: 980px) 980px, (min-width: 981px) and (max-width: 1280px) 1280px, (min-width: 1281px) 2560px, 100vw" /><p id="caption-attachment-104533" class="wp-caption-text">Figure 1. Agentic Architect, a framework for Computer Architecture Design Space Exploration and Optimization.</p></div>
<h2><b>Why the loop closes here</b></h2>
<p><span style="font-weight: 400;">Computer architecture has a structural advantage that is easy to take for granted: from the beginning, the field has organized itself around shared, quantitative empirical evaluation.</span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://www.spec.org/cpu2026/"><span style="font-weight: 400;"> SPEC</span></a><span style="font-weight: 400;">,</span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://www.cs.princeton.edu/techreports/2008/811.pdf"><span style="font-weight: 400;"> PARSEC</span></a><span style="font-weight: 400;">,</span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://www.cloudsuite.ch/"> <span style="font-weight: 400;">CloudSuite</span></a><span style="font-weight: 400;">,</span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://www.csl.cornell.edu/~delimitrou/papers/2019.asplos.microservices.pdf"><span style="font-weight: 400;"> DeathStarBench</span></a><span style="font-weight: 400;">, and a long list of others encode a community-wide agreement about what a &#8220;fair comparison&#8221; looks like. The metrics are equally well established: IPC, MPKI, miss rates, area, power, energy-delay product. Every paper in the field is, in effect, a measurement against an agreed instrument.</span></p>
<p><span style="font-weight: 400;">An agentic loop needs exactly this kind of discipline to close. The loop&#8217;s productivity is bounded by the cost and clarity of its fitness signal: how cheaply can a candidate be evaluated, and how reliably does the resulting score reflect the property we actually care about? In domains where evaluation is subjective, expensive, or contested, agentic exploration struggles. In computer architecture, the cycle-accurate simulator gives the loop reproducibility: controlled-environment evaluation against well-defined metrics. Production profiling, hardware performance counters, tracing, and system telemetry give it realism: behavior under load and access patterns that synthetic benchmarks cannot reproduce. The two together are what close the loop.</span></p>
<p><span style="font-weight: 400;">That has a practical consequence. It means the field does not have to invent its evaluation infrastructure to take advantage of agentic co-design; it has to </span><i><span style="font-weight: 400;">connect</span></i><span style="font-weight: 400;"> it. The benchmarks, the simulators, and the metric vocabulary are already in place. What is missing is the throughput and the integration: simulators that can serve hundreds of evaluations per study, fitness functions that compose IPC with area and power as primary terms, and training/evaluation splits that let us measure generalization instead of overfitting. We will return to this agenda below.</span></p>
<h2><b>When code is co-authored, what does &#8220;programmable&#8221; mean?</b></h2>
<p><span style="font-weight: 400;">Some of the architectural conservatism we aimed to maintain was justified, decades ago, by a single phrase: </span><i><span style="font-weight: 400;">but no one will program it</span></i><span style="font-weight: 400;">. The Cell processor&#8217;s programmer-managed SPEs and local stores are an example: an elegant design that proved very hard to program in practice.</span></p>
<p><span style="font-weight: 400;">The cost of programmability used to be borne almost entirely by humans. That is no longer true, and it changes the calculation.</span></p>
<p><span style="font-weight: 400;">In April 2025, Satya Nadella reported that</span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://techcrunch.com/2025/04/29/microsoft-ceo-says-up-to-30-of-the-companys-code-was-written-by-ai/"> <span style="font-weight: 400;">20% to 30% of Microsoft&#8217;s code</span></a><span style="font-weight: 400;"> was AI-generated, with internal acceptance rates rising monotonically. Google sits in a similar regime:</span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://fortune.com/2024/10/30/googles-code-ai-sundar-pichai/"> <span style="font-weight: 400;">a quarter in Q3 2024</span></a><span style="font-weight: 400;">, half by fall 2025, and</span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://blog.google/innovation-and-ai/infrastructure-and-cloud/google-cloud/cloud-next-2026-sundar-pichai/"> <span style="font-weight: 400;">75% of Google&#8217;s code by April 2026</span></a><span style="font-weight: 400;">, with Sundar Pichai describing the shift as &#8220;truly agentic workflows&#8221; in which engineers orchestrate fleets of AI agents rather than writing each line themselves.</span></p>
<p><span style="font-weight: 400;">These numbers describe authorship of characters, not accountability. But they should change how we evaluate the programmability constraint. When agents can routinely program across ISAs, generate platform-specific code paths, write test harnesses, and bridge unfamiliar interfaces given a clear specification, the cost on the programmer is no longer a sufficient veto on a hardware design choice. Designs that were dismissed because they imposed too high a cost on human programmers warrant a fresh look when most of that cost falls on agents instead.</span></p>
<p><span style="font-weight: 400;">Programmability still matters. Clarity, debuggability, verifiability, and predictable performance remain real properties humans need, and increasingly properties </span><i><span style="font-weight: 400;">agents</span></i><span style="font-weight: 400;"> need too. Abstractions still matter, perhaps more than ever. Deciding which to expose, which to hide, and which to make machine-checkable is now a question for the programming-languages and systems community alongside architects. But the most consequential lever may not be what we add; it may be what we remove. Many of the layers in today&#8217;s stack exist to hide hardware from human programmers and cost cycles and area to maintain. When agents absorb that complexity, the layers come off, and the performance and efficiency we have been paying to abstract away come back.</span></p>
<h2><b>A widening design space</b></h2>
<p><span style="font-weight: 400;">Reshaping the loop only matters if the space it has to cover is tractable. Increasingly, it isn&#8217;t.</span></p>
<p><span style="font-weight: 400;">A modern AI datacenter spans CPUs, GPUs, AI accelerators, a widening memory landscape (DRAM, CXL, HBM, HBF, SSD), and rack-scale integration with NVLink and optical interconnects. The software and hardware layers have not caught up: a single agentic query may dispatch dozens of model invocations across heterogeneous devices and tool calls on CPUs, all with different abstractions. </span><b>The operating system (OS), the layer that has historically reconciled such mismatches, must evolve at an unprecedented pace to keep up with growing hardware capabilities and software demands. </b><span style="font-weight: 400;">For example, it only has partial visibility into the GPU, despite its prominent role in AI workloads. Our recent work </span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://dl.acm.org/doi/10.1145/3731569.3764818"><span style="font-weight: 400;">LithOS</span></a><span style="font-weight: 400;"> has established a beachhead for OS-level control over GPUs, but extending that contract to coordinate the full heterogeneous stack is open. At every level of that stack, energy and power are hardening from secondary considerations into primary constraints.</span></p>
<p><span style="font-weight: 400;">Each of these pressures is, individually, a multi-year research program. </span><i><span style="font-weight: 400;">Together</span></i><span style="font-weight: 400;">, they describe a design space defined by heterogeneous compute, evolving memory hierarchies, rack-scale integration, software-level coordination, and workload regimes that did not exist five years ago. Covering this space by hand is increasingly difficult, even for a large team of architects.</span></p>
<p><span style="font-weight: 400;">That is the practical case for agentic co-design. The space is outgrowing human-only exploration, and the tools to cover it are finally here.</span></p>
<h2><b>A proof point</b></h2>
<p><span style="font-weight: 400;">In our recent work, we introduce the</span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://arxiv.org/abs/2604.25083"> <span style="font-weight: 400;">Agentic Architect</span></a><span style="font-weight: 400;">, an agentic framework for architecture design space exploration and optimization. We evaluate it across three of the most studied microarchitectural domains: cache replacement, data prefetching, and branch prediction. We chose them precisely </span><i><span style="font-weight: 400;">because</span></i><span style="font-weight: 400;"> they are mature. They have decades of literature, well-understood baselines, and limited remaining headroom; if the loop produces gains in these domains, the result is meaningful. The evolved cache replacement policy matched and slightly exceeded Mockingjay; the evolved prefetcher beat SOTA by 17%; the evolved branch predictor improved over Hashed Perceptron on workloads where branch behavior is the bottleneck.</span></p>
<div id="attachment_104534" style="width: 779px" class="wp-caption aligncenter"><img loading="lazy" decoding="async" aria-describedby="caption-attachment-104534" class="wp-image-104534 " src="https://www.sigarch.org/wp-content/uploads/2026/05/prefetch_scatter_final-3840x1566.png" alt="" width="769" height="314" /><p id="caption-attachment-104534" class="wp-caption-text">Figure 2. Storage versus performance for data prefetchers. The evolved prefetcher (87 KB) is Pareto-optimal: it delivers the highest geomean speedup over no prefetching at a smaller storage budget than the next-best design.</p></div>
<p><span style="font-weight: 400;">The more interesting result is what the loop discovered, and what it didn&#8217;t. The components in the evolved designs are almost entirely known techniques: stride engines and delta correlators for prefetching, reuse-distance predictors and signature tables for replacement, perceptron variants for branch prediction. None of these primitives is new.</span></p>
<p><span style="font-weight: 400;">What is new is the </span><i><span style="font-weight: 400;">coordination</span></i><span style="font-weight: 400;">. The evolved prefetcher continuously re-evaluates each predictive engine and throttles speculative ones under memory pressure. The evolved replacement policy arbitrates between three independent predictors based on their recent accuracy. The recurring structure across all three domains is the same: preserve the seed&#8217;s core, add orthogonal known features, integrate them through new coordination, and adapt at runtime. The novelty lies in the coordination. The loop refines the foundation; the architect still chooses it.</span></p>
<h2><b>The infrastructure needs to evolve</b></h2>
<p><span style="font-weight: 400;">If agentic co-design is going to do useful work across this design space, the bottleneck moves to infrastructure. The benchmarks and metrics are already there. What we need to build is throughput, multi-objective scoring, and cross-layer reach. The agenda is concrete:</span></p>
<ul>
<li style="font-weight: 400;"><b>New tools for agentic architecture design space exploration.</b><span style="font-weight: 400;"> Cycle-accurate simulators were built for human-paced experimentation; an agentic loop wants hundreds of evaluations per study, with storage, area, power, and timing as terms in the score rather than afterthoughts that disqualify the result later. We need simulators, search strategies, and metrics purpose-built for this regime: search loops that respect hardware constraints and balance exploration against exploitation, and composite metrics that combine performance, area cost, and generalization into signals the search can rank against.</span></li>
<li style="font-weight: 400;"><b>Cross-component and cross-layer co-evolution.</b><span style="font-weight: 400;"> Co-design across the OS/hardware boundary is now the norm rather than the exception. Taking virtual memory as an example, TLB design, page-table walkers, translation footprint in the caches, and huge-page promotion in the kernel are tightly coupled, and optimizing any one in isolation may capture only a fraction of the available improvement. RTL backends, full-system simulators, and formal verification each let the loop close around a different surface.</span></li>
<li style="font-weight: 400;"><b>Open source has to evolve.</b><span style="font-weight: 400;"> Releasing code is no longer enough. We need structured artifacts that span the full stack, from prompt, seed, and scoring function down to traces, simulator and system configurations, and where applicable RTL, packaged so an agent can clone a repo, re-run the search that produced a published result, and compare new candidates against the same baseline.</span></li>
</ul>
<p><span style="font-weight: 400;">The architecture and systems communities are uniquely positioned to drive that work.</span></p>
<h2><b>Renegotiating computer architecture and systems</b></h2>
<p><span style="font-weight: 400;">A stack co-authored by humans and agents needs renegotiation along the three axes of the old contract. Each is now reweighed against a new deliverable: </span><b>programmability</b><span style="font-weight: 400;"> for agents and humans alike, rather than humans alone.</span></p>
<p><i><span style="font-weight: 400;"><strong>Abstractions</strong>.</span></i><span style="font-weight: 400;"> Many layers exist precisely to hide hardware from human programmers, and they cost cycles and area to maintain. With agents absorbing that complexity, some of those layers can come off; performance and efficiency we have been paying to abstract away come back.</span></p>
<p><i><span style="font-weight: 400;"><strong>Interfaces</strong>.</span></i><span style="font-weight: 400;"> The boundary between hardware and software was drawn for human programmers. As agents become the primary author of low-level code, the interface that carries the contract forward needs redrawing: machine-checkable, composable, and accessible to tools rather than only to humans.</span></p>
<p><i><span style="font-weight: 400;"><strong>Transparency</strong>.</span></i><span style="font-weight: 400;"> The property that lets a programmer model the CPU in their head gives way to a stricter need: </span>explainability<span style="font-weight: 400;">. The architect must verify the result against intent, explain why it works, and check that it generalizes beyond the workloads it was trained on. None of these come for free; the field needs methods, metrics, and tooling that make them routine.</span></p>
<p><span style="font-weight: 400;">Leiserson and colleagues told us in 2020 that there was plenty of room at the Top of the computing stack. The half-decade since has been about confirming they were right; the next half-decade will be about whether we build the tools to actually live there. Agentic co-design, paired with learning embedded inside the system itself, is a strong candidate for addressing the &#8220;opportunistic, uneven, sporadic&#8221; character that delivered those gains so far.</span></p>
<p><span style="font-weight: 400;">Architecture is changing. The contract still holds, but the terms are up for negotiation. The next generation of the stack will be defined as much by what we remove as by what we add. The people best positioned to make those calls are the ones who understand both the hardware and the software. That is, by definition, our community.</span></p>
<p><b>Acknowledgments</b></p>
<p><span style="font-weight: 400;">Thanks to the Computer Architecture &amp; Operating System (</span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://www.cs.cmu.edu/~caos/"><span style="font-weight: 400;">CAOS</span></a><span style="font-weight: 400;">) group at Carnegie Mellon and to Prof. Alex Daglis and Prof. Todd Mowry for feedback on this post.</span></p>
<p><b>About the Author</b></p>
<p><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://www.cs.cmu.edu/~dskarlat/"><span style="font-weight: 400;">Dimitrios Skarlatos</span></a><span style="font-weight: 400;"> is an assistant professor in the Computer Science Department at Carnegie Mellon University. His research bridges computer architecture and operating systems with a focus on AI datacenter efficiency, privacy, and scalability. His work has been deployed in production datacenters and upstreamed into the Linux kernel. He has received the IEEE CS TCCA Young Computer Architect Award, the NSF CAREER Award, the Intel Rising Star Award, a Linux Foundation Faculty Award, an ISCA Best Paper Award, two ASPLOS Best Paper Awards, a CACM Research Highlight, four IEEE MICRO Top Picks, the joint ACM SIGARCH &amp; IEEE CS TCCA Outstanding Dissertation Award, the David J. Kuck Outstanding PhD Thesis Award, and over a dozen industry faculty awards from Amazon, AMD, Intel, Meta, Oracle, and VMware. His recent work led to the founding of </span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://lithosai.com/"><span style="font-weight: 400;">LithosAI</span></a><span style="font-weight: 400;">.</span></p>
<p>&nbsp;</p>
<p class="disclaim"><strong>Disclaimer:</strong> <em>These posts are written by individual contributors to share their thoughts on the Computer Architecture Today blog for the benefit of the community. Any views or opinions represented in this blog are personal, belong solely to the blog author and do not represent those of ACM SIGARCH or its parent organization, ACM.</em></p><Img align="left" border="0" height="1" width="1" alt="" style="border:0;float:left;margin:0;padding:0;width:1px!important;height:1px!important;" hspace="0" src="https://feeds.feedblitz.com/~/i/956665028/0/sigarch-cat">
]]>
</content:encoded>
			<wfw:commentRss>https://feeds.feedblitz.com/~/956665028/0/sigarch-cat/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
	<post-id xmlns="com-wordpress:feed-additions:1">104529</post-id></item>
<item>
<feedburner:origLink>https://www.sigarch.org/from-control-to-data-to-value-a-third-axis-of-parallelism/</feedburner:origLink>
		<title>From Control to Data to Value: A Third Axis of Parallelism</title>
		<link>https://feeds.feedblitz.com/~/955857566/0/sigarch-cat/</link>
		<comments>https://feeds.feedblitz.com/~/955857566/0/sigarch-cat/#respond</comments>
		<pubDate>Wed, 13 May 2026 15:00:29 +0000</pubDate>
		<dc:creator><![CDATA[Di Wu, Zhewen Pan, Joshua San Miguel]]></dc:creator>
		<category><![CDATA[ACM SIGARCH]]></category>
		<category><![CDATA[AI hardware]]></category>
		<category><![CDATA[Parallelism]]></category>
		<guid isPermaLink="false">https://www.sigarch.org/?p=104410</guid>
		<description><![CDATA[<div><img width="300" xheight="176" src="https://www.sigarch.org/wp-content/uploads/2026/05/vlp-sigarch-blog-3-300x176.png" class="attachment-medium size-medium wp-post-image" alt="" style="max-width:100% !important;height:auto !important;margin-bottom:15px;margin-left:15px;float:right;"  loading="lazy" /></div>TL;DR: The history of parallel computing is a history of shifting what we put at the center of the computer. The first axis, control-level parallelism (CLP), is control-centric and schedules around the program counter: it gave us the high-performance computing (HPC) era. The second axis, data-level parallelism (DLP), is data-centric and schedules around tensors: it [&#8230;]]]>
</description>
				<content:encoded><![CDATA[<div><img width="300" height="176" src="https://www.sigarch.org/wp-content/uploads/2026/05/vlp-sigarch-blog-3-300x176.png" class="attachment-medium size-medium wp-post-image" alt="" style="margin-bottom:15px;margin-left:15px;float:right;" decoding="async" loading="lazy" /></div><p><strong>TL;DR:</strong> <span style="font-weight: 400;">The history of parallel computing is a history of shifting what we put at the center of the computer. The first axis, control-level parallelism (CLP), is control-centric and schedules around the program counter: it gave us the high-performance computing (HPC) era. The second axis, data-level parallelism (DLP), is data-centric and schedules around tensors: it gave us the artificial intelligence (AI) era. A third axis is now emerging: </span><i><span style="font-weight: 400;">value-level parallelism (VLP)</span></i><span style="font-weight: 400;">, where narrow data bitwidth exposes a small number of unique values and lets the architecture deduplicate redundant computation. Two recent works, Carat (ASPLOS &#8217;24) and Mugi (ASPLOS &#8217;26), make the case concretely: VLP eliminates redundant computation in both linear and nonlinear operations on AI workloads. This article argues that VLP is not a point of optimization but the beginning of a </span><i><span style="font-weight: 400;">value-centric computing</span></i><span style="font-weight: 400;"> paradigm, one that is crucial for addressing the escalating energy demands of next-generation intelligent systems.</span></p>
<p>&nbsp;</p>
<h1><strong>Traditional Parallel Computing</strong></h1>
<h3><strong>The First Axis: Control-Level Parallelism</strong></h3>
<p><span style="font-weight: 400;">The HPC era saw a diverse set of workloads. The metric of success made the goal explicit: instructions per cycle (IPC) normalizes performance to </span><i><span style="font-weight: 400;">how fast instructions are consumed</span></i><span style="font-weight: 400;">, not to what data the instructions are operating on. Therefore, computer architecture in the HPC era was control-centric. Michael Flynn formalized the design space in 1966 with his taxonomy </span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://dl.acm.org/doi/10.5555/333067.333226"><span style="font-weight: 400;">[1]</span></a><span style="font-weight: 400;">: SISD, SIMD, MISD, MIMD, outlining the orthogonality of instruction and data. For decades, this was the right framing: transistors were scarce, control logic was expensive, and the most valuable thing an architect could do was to issue one more instruction per cycle. </span></p>
<p><span style="font-weight: 400;">The actual implementation to exploit the CLP to execute multiple instructions in parallel arose from the independence between instructions. We enjoyed the technology evolution from pipelining, branch prediction for SISD, superscalar issue, out-of-order execution, simultaneous multithreading, chip multiprocessors for MIMD and beyond.</span></p>
<h3><strong>The Second Axis: Data-Level Parallelism</strong></h3>
<p><span style="font-weight: 400;">Entering the AI era, powered by large language models (LLMs), transistors became plentiful but the memory bandwidth became scarce due to the large data volume in AI tensors. Consequently, we hit the memory wall and turned to data-centric architectures, expanding more along the data dimension in Flynn&#8217;s taxonomy. Success is now measured by </span><i><span style="font-weight: 400;">how well we apply one operation to many data elements while feeding them efficiently from memory</span></i><span style="font-weight: 400;"> (e.g., throughput, goodput, and arithmetic intensity). TPUs with systolic arrays and GPUs with tensor cores are renowned examples to exploit the rich DLP opportunities from high-dimensional tensors.</span></p>
<p><span style="font-weight: 400;">Diving deeper, we see that compute arrays, i.e., dataflow architecture, are becoming the first class citizens. Dataflow architecture follows the philosophy of </span><i><span style="font-weight: 400;">letting data drive control</span></i><span style="font-weight: 400;">, with early works from MIT </span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://dl.acm.org/doi/10.1145/642089.642111"><span style="font-weight: 400;">[2]</span></a> <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://dl.acm.org/doi/10.1109/12.48862"><span style="font-weight: 400;">[3]</span></a><span style="font-weight: 400;">. With regular compute and memory patterns in AI tensors, dataflow architecture builds massively parallel compute arrays to maximize the computational density and minimize the control overhead. </span></p>
<p><span style="font-weight: 400;">Another line of research to exploit DLP for AI workloads targets the von Neumann bottleneck, envisioned by John Backus as early as 1978 </span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://dl.acm.org/doi/10.1145/359576.359579"><span style="font-weight: 400;">[4]</span></a><span style="font-weight: 400;">. These solutions move the computation closer to the memory (e.g., in/near-memory/storage processing) by attaching additional compute logic next to the memory blocks, unlocking massive DLP on the already wide enough memory blocks.</span></p>
<p>&nbsp;</p>
<h1><strong>The Third Axis: Value-Level Parallelism</strong></h1>
<p><span style="font-weight: 400;">Though Flynn’s taxonomy, with dimensions of </span><i><span style="font-weight: 400;">instructions</span></i><span style="font-weight: 400;"> and </span><i><span style="font-weight: 400;">data</span></i><span style="font-weight: 400;">,</span> <span style="font-weight: 400;">has been followed for decades, there are untouched landscapes. While CLP and DLP focus on concurrency and parallelism, neither asks the next question about the </span><i><span style="font-weight: 400;">content</span></i><span style="font-weight: 400;"> of the data: </span><i><span style="font-weight: 400;">can the patterns in data values benefit the computation efficiency? </span></i><span style="font-weight: 400;">VLP is value-centric and the third axis in this regard, i.e., it targets </span><i><span style="font-weight: 400;">computational redundancy </span></i><span style="font-weight: 400;">inherent to the data patterns of workloads. It recognizes that when identical values flow through a pipeline, the arithmetic becomes deterministic and, therefore, avoidable. Consequently, we move from executing every instruction and data to computing only each unique data value.</span></p>
<h3><strong>Origins for GEMM</strong></h3>
<p><b>Carat (Pan, San Miguel, Wu — ASPLOS &#8217;24)</b> <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://dl.acm.org/doi/10.1145/3620665.3640364"><span style="font-weight: 400;">[5]</span></a><span style="font-weight: 400;"> is the paper that coined VLP and materialized it in hardware architecture. The insight is simple. As deep learning inference moves to larger batches and lower precisions (e.g., FP8 being the de-facto data format in DeepSeek v3), the number of </span><i><span style="font-weight: 400;">unique</span></i><span style="font-weight: 400;"> values shrinks rapidly while the frequency of each grows. Here, we give an example. For a scalar-vector multiplication for an arbitrary scale weight </span><i><span style="font-weight: 400;">w</span></i><span style="font-weight: 400;"> and 1k UINT4 inputs, conventional hardware would compute 1k multiplications for the weight </span><i><span style="font-weight: 400;">w </span></i><span style="font-weight: 400;">and each UINT4 vector element. Looking closely, there are only 16 unique products, i.e., 0 x </span><i><span style="font-weight: 400;">w</span></i><span style="font-weight: 400;">, 1 x </span><i><span style="font-weight: 400;">w, </span></i><span style="font-weight: 400;">2 x </span><i><span style="font-weight: 400;">w, …, </span></i><span style="font-weight: 400;">15 x </span><i><span style="font-weight: 400;">w.</span></i><span style="font-weight: 400;"> Thus, conventional hardware would compute 1k/16=64 times more than needed.</span></p>
<p><img loading="lazy" decoding="async" class="alignnone wp-image-104478" src="https://www.sigarch.org/wp-content/uploads/2026/05/vlp-sigarch-blog-1-scaled.png" alt="" width="515" height="289" /></p>
<p><span style="font-weight: 400;">Figure 1. Overview of VLP for scalar-vector multiplication.</span></p>
<p><span style="font-weight: 400;">Figure 1 above outlines how VLP is constructed in Carat. In Figure 1 (a), VLP consists of value reuse, which accumulates the weight </span><i><span style="font-weight: 400;">w</span></i><span style="font-weight: 400;"> over time, each accumulation result, called a partial product, is used to compute the next partial product. Then each input just </span><i><span style="font-weight: 400;">subscribes </span></i><span style="font-weight: 400;">to the proper partial product as the correct output. To materialize the subscription, we leverage temporal coding, often seen in the brain, which generates a spike at the cycle indexed by the data value. For example, a data valued 8 will generate a spike at cycle 8, as shown in Figure 1 (b). Therefore, there exists a temporal correspondence between the spike and the accumulated partial product. Each input subscribes to their correct output in parallel, giving the rise to value-level parallelism. The scheduling unit is no longer the instruction or the array element; it is the unique product value, made available to many input consumers via temporal coding. Given this formulation, we see that lower precision produces fewer unique values, while larger batches create more inputs to share the unique values.</span></p>
<h3><strong>Generalizing Beyond</strong></h3>
<p><span style="font-weight: 400;">So far, VLP in Carat targets GEMM optimization for large-batch, low-precision, symmetric-format use cases. However, these assumptions may no longer hold in more recent LLM workloads: the batch size is small (e.g., 8~16) to ensure real-time response, the data formats are asymmetric (e.g., INT4-FP16) to minimize the memory footprint of the weight and KV cache, and the nonlinear operations are heavy and complicated (e.g., softmax, GELU, SiLU) to ensure high accuracy.</span></p>
<p><b>Mugi (Price, Vellaisamy, Shen, Wu — ASPLOS &#8217;26)</b> <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://dl.acm.org/doi/10.1145/3779212.3790189"><span style="font-weight: 400;">[6]</span></a><span style="font-weight: 400;"> is a follow-up work that closes the gap and generalizes VLP for both linear and nonlinear operations for broader AI workloads. Figure 2 shows VLP for elementwise nonlinear operations that can be done through a four-phase pipeline on input floating-point numbers with a sign (S), mantissa (M) and exponent (E). The first phase is input approximation, which converts wider inputs to narrower bits without sacrificing the LLM accuracy too much. The inputs to nonlinear operations are always in higher precision (e.g., BF16, FP16, or FP32) in AI workloads than model weights. The approximation to narrower bits ensures a shorter temporal signal for high throughput. The second phase is value reuse with opportunities from large GEMM shapes in LLMs. Unlike value reuse in Carat accumulating the partial product, value reuse in Mugi loads the precomputed nonlinear results, where higher accuracy is allocated to more critical inputs. The third phase performs temporal subscription on the mantissa bits (M) of the inputs. For each input, the selected output corresponds to the nonlinear results with the same mantissa but different exponents (E). Finally, the fourth phase performs temporal subscription on the exponent bits of the inputs. For each input, the selected output corresponds to the correct mantissa and exponent.</span></p>
<p><img loading="lazy" decoding="async" class="alignnone wp-image-104477" src="https://www.sigarch.org/wp-content/uploads/2026/05/vlp-sigarch-blog-2-scaled.png" alt="" width="726" height="173" /></p>
<p><span style="font-weight: 400;">Figure 2. Overview of VLP for elementwise nonlinear operations.</span></p>
<p><span style="font-weight: 400;">Mugi essentially cascades VLP to construct high dimensionality in the value space to compute nonlinear operations. The resulting architecture unifies the datapath for both linear and nonlinear operations, leading to savings in silicon area.</span></p>
<h3><strong>What Makes VLP Different</strong></h3>
<p><span style="font-weight: 400;">Conventional architectures pay arithmetic costs to process every instruction and data element while advanced techniques (e.g., memoization, tabulation) leverage computation reuse and pay memory costs to store past results and refer back to them when needed. VLP pays for </span><i><span style="font-weight: 400;">neither</span></i><span style="font-weight: 400;">: it minimizes both arithmetic and memory access, instead spending its silicon budget on </span><i><span style="font-weight: 400;">value delivery</span></i><span style="font-weight: 400;">, i.e., the network and temporal converters that route unique results to many input consumers in parallel. The type of computation fundamentally changes to something new: a form of </span><i><span style="font-weight: 400;">temporal subscription</span></i><span style="font-weight: 400;">.</span></p>
<p>&nbsp;</p>
<h1><strong>Future Opportunities and Open Questions</strong></h1>
<p><span style="font-weight: 400;">In its current form, VLP relies on temporal coding, which is actually inspired from how the brain works. Though the community has focused predominantly on deep learning, VLP opens up a different research direction: </span><i><span style="font-weight: 400;">what are potential synergies between neuromorphic and classical computing in computer architecture</span></i><span style="font-weight: 400;">? Looking beyond, VLP raises several questions to answer in the AI era. </span></p>
<ul>
<li><i><span style="font-weight: 400;">Where does VLP stop paying off? </span></i><span style="font-weight: 400;">Though Carat and Mugi work for varying batch sizes, both of them now are designed for low precision to create more opportunities for value reuse. There is presumably a design space with high precision and low value redundancy. It is essential to understand the mechanism to exploit VLP for such scenarios and quantify the potential gain.</span></li>
<li><i>Is VLP an ISA-level concept or an accelerator-level one? </i>Both Carat and Mugi are accelerator designs. A real-world question is whether VLP can inform CPU and GPU microarchitecture. What would a VLP-based tensor instruction look like? Could it be a drop-in replacement of tensor cores with better efficiency?</li>
<li><i>What should the software stack look like?</i> VLP for nonlinear involves approximation, which naturally introduces inaccuracy to the task. This falls back to the question of approximate computing, but in the context of new hardware primitives. We probably shall build a co-design framework to deploy VLP under approximation errors.</li>
<li><i>Can VLP live with sparsity? </i>Sparse computation has been a major optimization since the start of deep learning, and more opportunities are emerging from weight, KV-cache spanning across the bit level, value level, block level and even request level. It is meaningful to study how VLP synergizes with such use cases.</li>
<li><i>How does VLP interact with memory? </i>Despite efforts in computation, memory stays at the core of AI. A natural question is whether we can optimize the memory system with VLP, or more broadly, value-centric computing. There have been associative memory-based AI accelerators, and whether there could be VLP-based alternatives?</li>
</ul>
<p>&nbsp;</p>
<h1><strong>Final Thoughts</strong></h1>
<p><span style="font-weight: 400;">Architecture research is all about how to compute faster and more efficiently. </span><i><span style="font-weight: 400;">Control-level parallelism</span></i><span style="font-weight: 400;"> has let us argue about IPC and pipelines for thirty years. </span><i><span style="font-weight: 400;">Data-level parallelism</span></i><span style="font-weight: 400;"> has let us argue about FLOPS and dataflow for fifteen. </span><i><span style="font-weight: 400;">Value-level parallelism</span></i><span style="font-weight: 400;">, the third axis, now shows promise for emerging AI workloads (thanks to Carat and Mugi) and paves the way for exciting synergies between neuromorphic and classical computing. Here&#8217;s hoping one day computer architects see the </span><i><span style="font-weight: 400;">value</span></i><span style="font-weight: 400;"> in it.</span></p>
<p>&nbsp;</p>
<h3><span style="font-weight: 400;"><strong>About the authors:</strong> </span></h3>
<p><i><span style="font-weight: 400;"><strong>Di Wu</strong> is an assistant professor at the University of Central Florida. </span></i></p>
<p><i><span style="font-weight: 400;"><strong>Zhewen Pan</strong> is a PhD candidate at the University of Wisconsin–Madison. </span></i></p>
<p><i><span style="font-weight: 400;"><strong>Joshua San Miguel</strong> is an associate professor at the University of Wisconsin–Madison.</span></i></p>
<p class="disclaim"><strong>Disclaimer:</strong> <em>These posts are written by individual contributors to share their thoughts on the Computer Architecture Today blog for the benefit of the community. Any views or opinions represented in this blog are personal, belong solely to the blog author and do not represent those of ACM SIGARCH or its parent organization, ACM.</em></p><Img align="left" border="0" height="1" width="1" alt="" style="border:0;float:left;margin:0;padding:0;width:1px!important;height:1px!important;" hspace="0" src="https://feeds.feedblitz.com/~/i/955857566/0/sigarch-cat">
]]>
</content:encoded>
			<wfw:commentRss>https://feeds.feedblitz.com/~/955857566/0/sigarch-cat/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
	<post-id xmlns="com-wordpress:feed-additions:1">104410</post-id></item>
<item>
<feedburner:origLink>https://www.sigarch.org/how-ai-will-reshape-computer-systems-by-2035-a-jeffersonian-dinner-in-san-francisco-about-our-10000x-future/</feedburner:origLink>
		<title>How AI Will Reshape Computer Systems by 2035: A Jeffersonian Dinner in San Francisco about Our 10,000x Future</title>
		<link>https://feeds.feedblitz.com/~/955221287/0/sigarch-cat/</link>
		<comments>https://feeds.feedblitz.com/~/955221287/0/sigarch-cat/#respond</comments>
		<pubDate>Mon, 04 May 2026 14:00:06 +0000</pubDate>
		<dc:creator><![CDATA[Helen Wright, Jeff Dean, Mark D. Hill, and Dave Patterson]]></dc:creator>
		<category><![CDATA[ACM SIGARCH]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[Computer Systems]]></category>
		<guid isPermaLink="false">https://www.sigarch.org/?p=103841</guid>
		<description><![CDATA[<div><img width="300" xheight="239" src="https://www.sigarch.org/wp-content/uploads/2026/04/Screenshot-2026-04-30-at-3.55.21-PM-300x239.png" class="attachment-medium size-medium wp-post-image" alt="" style="max-width:100% !important;height:auto !important;margin-bottom:15px;margin-left:15px;float:right;"  loading="lazy" /></div>Editor&#8217;s Note: this post is a republication of CRA-I post available at: https://cra.org/industry/2026/04/27/how-ai-will-reshape-computer-systems-by-2035-a-jeffersonian-dinner-in-san-francisco-about-our-10000x-future/ CRA-Industry (CRA-I) recently continued its series of intimate Industry Salon Gatherings, bringing together leaders to discuss the long-term trajectory of our field. Our latest session, organized by Mark D. Hill (University of Wisconsin-Madison &#38; CRA) and CRA-I, took place on April 16, [&#8230;]]]>
</description>
				<content:encoded><![CDATA[<div><img width="300" height="239" src="https://www.sigarch.org/wp-content/uploads/2026/04/Screenshot-2026-04-30-at-3.55.21-PM-300x239.png" class="attachment-medium size-medium wp-post-image" alt="" style="margin-bottom:15px;margin-left:15px;float:right;" decoding="async" loading="lazy" /></div><p><em>Editor&#8217;s Note: this post is a republication of CRA-I post available at: <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://cra.org/industry/2026/04/27/how-ai-will-reshape-computer-systems-by-2035-a-jeffersonian-dinner-in-san-francisco-about-our-10000x-future/">https://cra.org/industry/2026/04/27/how-ai-will-reshape-computer-systems-by-2035-a-jeffersonian-dinner-in-san-francisco-about-our-10000x-future/</a></em></p>
<p><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://cra.org/industry/">CRA-Industry (CRA-I)</a> recently continued its series of intimate Industry Salon Gatherings, bringing together leaders to discuss the long-term trajectory of our field. Our latest session, organized by Mark D. Hill (University of Wisconsin-Madison &amp; CRA) and CRA-I, took place on April 16, 2026 at the historic <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://www.uclubsf.org/">University Club of San Francisco</a> and was sponsored by <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://www.laude.org/">Laude Institute</a>, which had just announced an ambitious and exciting slate of new <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://www.laude.org/moonshots">research “AI moonshot” awards</a>.</p>
<p>The evening featured a high-level conversation among 20 participants, co-hosted by <b>Dave Patterson </b>(UC Berkeley/Google) and <b>Jeff Dean </b>(Google AI)<b>.</b> The Salon tackled a fundamental question: “<a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://cra.org/industry/events/what-will-computer-systems-look-like-in-2035-and-how-will-they-be-designed/">What will computer systems look like in 2035, and how will they be designed?”</a> The participants were researchers and leaders from West Coast academic institutions and technology companies, big and small. As the table below shows, the group was diverse in multiple dimensions, including career stage, with expertise centering on computer architecture but extending to software systems, AI models, and design methodology and tools.</p>
<p>Following the <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://www.bouteco.co/sustainable-luxury-travel/2022/4/27/whats-a-jeffersonian-dinner-earth-day">“Jeffersonian” dinner</a> format, the evening was designed to forge deep connections and brainstorm visionary ideas. To ensure a frank and pre-competitive dialogue, the event operated under the <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://www.chathamhouse.org/about-us/chatham-house-rule">Chatham House Rule</a>, allowing for candid exchange while protecting the anonymity of the participants’ specific contributions.</p>
<p>The three-hour (!) conversation explored the intersection of architecture, software, AI, and future design methodologies. Here we highlight some key observations and conjectures made by Salon participants:</p>
<ul>
<li>We are in the midst of an AI revolution that appears to be even more impactful than the introduction of microprocessors, PCs, Internet, or smartphones.</li>
<li>To drive future change, we must focus on metrics such as improving “intelligence” per Watt for efficiency and more AI tokens processed at fixed user-perceived latency for more “intelligence.”</li>
<li>One lively topic was centered on the future of interfaces and abstractions that have been essential for humans to build complex systems. We explored two related questions: whether abstractions will continue to matter in the AI era, and, if they do, whether those abstractions must remain human-interpretable, allowing human/AI teams to advance the field together. The general view was that abstractions will continue to play an important role, not only for humans, but also helping AI systems to reason and coordinate. However, for communication and reasoning among agents, these abstractions need not be interpretable to humans, and may evolve beyond human comprehension. At the same time, there remains a need for human-interpretable abstractions to enable oversight, intervention and guidance by humans.</li>
<li>We conjecture that in five years 10,000x more AI inference will be done worldwide, with these gains hypothetically coming from multiplicative progress of  50x in AI algorithms, 50x in system/hardware optimization/specialization, and 4x from further data center growth.</li>
<li>We expect 50x AI algorithm progress for three reasons. First, Transformers have sparked a rapid series of AI inventions since the <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://en.wikipedia.org/wiki/Attention_Is_All_You_Need">original paper in 2017</a>, and the area is still ripe. There has been a clear trend in increased data and compute efficiency in training large models. Second, there is an existence proof that we can do much better since humans learning from birth to early adulthood are able to use 1,000x less input data than today’s largest ML models, suggesting that much more data-efficient learning algorithms are still possible.Third, deep analysis to determine what aspects of current AI technology are necessary for current levels of intelligence may reveal much more efficient methods for obtaining the same or better results.</li>
<li>We expect 50x system/hardware progress from two trends. First, increased hardware specialization will lead to major improvements in efficiency. Second, AI to automate hardware design will enable much faster and lower-cost creation of this specialized hardware. AI is already greatly impacting the system/hardware design process. We expect many design flows to be accelerated or altered. For example, can Large Language Models iterate with improving formal tools and specification methods to “hill climb” design spaces, and can formal verification be accelerated by customized hardware? Moreover, AI is having a fundamental impact on software developers. We conjecture that future developers will write little code directly. Rather, they will manage teams of AI agents. And this in turn could dramatically accelerate software development for specialized hardware. How do we prepare students and professionals for this world?</li>
<li>We expect a substantial increase in global data center capacity (perhaps 4x over the next five years), but recognize that non-technical forces are at play here. Expansion will be relatively larger if companies focus on scale out more than the above innovation opportunities. However, expansion could be much less due to community “techlash.”</li>
<li>We discussed energy trends, including that solar was now significantly cheaper than other energy sources for new capacity, and that battery prices continued to fall, making solar+batteries a viable way of powering new datacenters and other energy needs.  We also discussed other new sources of clean energy that are not yet commercially viable but might be in the next five years, such as fusion.</li>
<li>Finally, acknowledging “tech-lash” brings us to opportunities and challenges that AI brings to society. <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://newrepublic.com/article/209163/ai-industry-discovering-public-backlash">Much of the public views AI as a potential disruption in their lives, fearing the negative more than embracing the positive</a>. They may currently be correct. It is our job to ensure that we develop and encourage the societally- beneficial aspects of AI, and we discussed education and healthcare as two domains with considerable early positive benefits and enormous further potential.  We also recognized negative aspects of AI usage in areas like easing the creation of misinformation and cyberattacks, and many expressed concern about society’s capacity to absorb substantial job disruption in short time scales.  We conjecture that within the next few years, AI policy will be a major election factor. How do we make a public AI debate substantive? How do we ensure policymakers can make informed decisions about AI, so that we can sensibly regulate some of the negative consequences of AI without stifling the positive uses? As AI changes our world, how does that change university teaching, life-long learning, and job retraining? How does it alter university and industry research? Three hours are insufficient to answer these societal questions.</li>
</ul>
<p>In their Turing Award lecture in 2018, <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://cacm.acm.org/research/a-new-golden-age-for-computer-architecture/">Hennessy and Patterson asserted that we were beginning a Golden Age for computer architecture</a>. The subsequent decade has shown them to be prescient.</p>
<p>This San Francisco gathering reinforced the value of bringing industry and academia together to look beyond the immediate product cycle. As noted by the organizers, the insights gained here will help shape the CRA roadmap for supporting the computing research ecosystem over the next decade.</p>
<p>As we went around the table to get everyone’s last comments, many mentioned how much they enjoyed the conversation and would love to do it again. Impressively, four of the invited leaders had to fly to San Francisco for the event and everyone stayed until the official end of the three-plus-hour reception and dinner. The feedback was overwhelmingly positive. One participant noted that it was “valuable to hear directly from the leaders of our community [about] their perspectives on the role of AI—both in how we design computers and the projected compute demand driving that design.” Another shared that the evening was “tremendously valuable for my own thinking on where the computing industry is and should be going, and for sharing thoughts to hopefully influence how others view technology advancement and its societal implications.”</p>
<p>If you are interested in participating in or hosting future CRA-I Salon Gatherings, please sign up for our mailing list or contact Helen Wright (<a href="mailto:hwright@cra.org">hwright@cra.org</a>).</p>
<p><b>Salon Participants</b></p>
<table>
<tbody>
<tr>
<td><b>First Name</b></td>
<td><b>Last Name</b></td>
<td><b>Affiliation</b></td>
</tr>
<tr>
<td>Doug</td>
<td>Burger</td>
<td>Microsoft Research</td>
</tr>
<tr>
<td>Jason</td>
<td>Cong</td>
<td>UCLA</td>
</tr>
<tr>
<td>Jeff</td>
<td>Dean</td>
<td>Google</td>
</tr>
<tr>
<td>Chris</td>
<td>Fletcher</td>
<td>Berkeley</td>
</tr>
<tr>
<td>Anna</td>
<td>Goldie</td>
<td>Ricursive Intelligence</td>
</tr>
<tr>
<td>Peter</td>
<td>Harsha</td>
<td>CRA</td>
</tr>
<tr>
<td>Mark</td>
<td>Hill</td>
<td>CRA/University of Wisconsin–Madison</td>
</tr>
<tr>
<td>Andy</td>
<td>Konwinski</td>
<td>Laude Institute</td>
</tr>
<tr>
<td>Alex</td>
<td>Ksendzovsky</td>
<td>The Biological Computing Co</td>
</tr>
<tr>
<td>Azalia</td>
<td>Mirhoseini</td>
<td>Stanford/Ricursive Intelligence</td>
</tr>
<tr>
<td>Dave</td>
<td>Patterson</td>
<td>Berkeley/Google</td>
</tr>
<tr>
<td>Chris</td>
<td>Ramming</td>
<td>CRA-Industry</td>
</tr>
<tr>
<td>Sophia</td>
<td>Shao</td>
<td>Berkeley</td>
</tr>
<tr>
<td>Ben</td>
<td>Spector</td>
<td>Flapping Airplanes</td>
</tr>
<tr>
<td>Ion</td>
<td>Stoica</td>
<td>Berkeley</td>
</tr>
<tr>
<td>Caroline</td>
<td>Trippel</td>
<td>Stanford</td>
</tr>
<tr>
<td>Natalia</td>
<td>Vassilieva</td>
<td>Cerebras Systems</td>
</tr>
<tr>
<td>Ralph</td>
<td>Wittig</td>
<td>AMD</td>
</tr>
<tr>
<td>Helen</td>
<td>Wright</td>
<td>CRA</td>
</tr>
<tr>
<td>Carole-Jean</td>
<td>Wu</td>
<td>Meta</td>
</tr>
</tbody>
</table>
<p>&nbsp;</p>
<p><strong>About the Authors:</strong></p>
<div><strong>Helen Wright</strong> is the Manager of CRA-Industry (CRA-I), a committee of the Computing Research Association. She bridges the gap between academia and industry, convening leadership to tackle &#8220;grand challenges&#8221; like workforce, socially responsible AI, and partnerships. Helen leads a Council and Steering Committee of over 20 industry pioneers to drive these initiatives forward. Previously, she was a Science Education Analyst at the NSF. She holds graduate and undergraduate degrees from the University of Virginia.</div>
<div></div>
<div><strong>Jeff Dean</strong> joined Google in 1999 and is currently Google’s Chief Scientist, where he co-leads the Gemini effort. His areas of focus include machine learning and AI, computer systems, and AI applications. He has worked on Google Search, Google News, Google Translate, Google’s advertising systems, MapReduce, BigTable, Spanner, TensorFlow, Pathways, and Gemini. In 2011, he co-founded the Google Brain project. He received a B.S. in CS and economics from the University of Minnesota in 1990 and a Ph.D. in CS from the University of Washington in 1996. He is a member of the U.S. National Academy of Engineering (2009), and is a Fellow of the ACM and the AAAS. He is a recipient of the ACM Prize in Computing, and the IEEE John von Neumann medal.</div>
<div></div>
<div><strong>Mark D. Hill</strong> is the Gene M. Amdahl and John P. Morgridge Professor Emeritus of Computer Sciences at the University of Wisconsin-Madison (<a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~www.cs.wisc.edu/~markhill">http://www.cs.wisc.edu/~markhill</a>), following his 1988-2020 service in CS and ECE. His research interests include parallel-computer system design, memory system design, and computer simulation. Hill&#8217;s work is highly collaborative with over 170 co-authors. He received the 2019 Eckert-Mauchly Award and is a fellow of AAAS, ACM, and IEEE. He serves on Computing Research Association (CRA) Board of Directors that is sponsoring this event. Hill was also Partner Hardware Architect at Microsoft (2020-24) where he led some software-hardware pathfinding for Azure.</div>
<div></div>
<div><strong>David Patterson</strong> retired after 40 years as an EECS professor at UC Berkeley before joining Google in 2016 as a Distinguished Engineer. He is probably best known for the book Computer Architecture: A Quantitative Approach and for the Berkeley RISC (Reduced Instruction Set Computer), RAID (Redundant Array of Inexpensive Disks) and NOW (Network of Workstations) projects. He and his co-author John Hennessy shared the 2017 ACM A.M Turing Award (the “Nobel Prize of Computing”) and the 2022 NAE Charles Stark Draper Prize for Engineering (a “Nobel Prize of Engineering”).</div>
<p class="disclaim"><strong>Disclaimer:</strong> <em>These posts are written by individual contributors to share their thoughts on the Computer Architecture Today blog for the benefit of the community. Any views or opinions represented in this blog are personal, belong solely to the blog author and do not represent those of ACM SIGARCH or its parent organization, ACM.</em></p><Img align="left" border="0" height="1" width="1" alt="" style="border:0;float:left;margin:0;padding:0;width:1px!important;height:1px!important;" hspace="0" src="https://feeds.feedblitz.com/~/i/955221287/0/sigarch-cat">
]]>
</content:encoded>
			<wfw:commentRss>https://feeds.feedblitz.com/~/955221287/0/sigarch-cat/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
	<post-id xmlns="com-wordpress:feed-additions:1">103841</post-id></item>
<item>
<feedburner:origLink>https://www.sigarch.org/an-overview-of-the-fourth-data-prefetching-championship-part-2/</feedburner:origLink>
		<title>Fourth Data Prefetching Championship: Part 2</title>
		<link>https://feeds.feedblitz.com/~/954807320/0/sigarch-cat/</link>
		<comments>https://feeds.feedblitz.com/~/954807320/0/sigarch-cat/#respond</comments>
		<pubDate>Wed, 29 Apr 2026 14:00:03 +0000</pubDate>
		<dc:creator><![CDATA[Digvijay Singh]]></dc:creator>
		<category><![CDATA[ACM SIGARCH]]></category>
		<category><![CDATA[Data Prefetcher]]></category>
		<category><![CDATA[Memory Wall]]></category>
		<guid isPermaLink="false">https://www.sigarch.org/?p=103014</guid>
		<description><![CDATA[<div><img width="300" xheight="187" src="https://www.sigarch.org/wp-content/uploads/2026/04/DPC-4-Part-1-300x187.png" class="attachment-medium size-medium wp-post-image" alt="DPC-4 Concept Art (Indigo)" style="max-width:100% !important;height:auto !important;margin-bottom:15px;margin-left:15px;float:right;"  loading="lazy" /></div>This article continues (and concludes) the discussion on the proceedings of DPC-4, covering the remaining four contestants and a summary of the trends observed in all eight prefetchers presented in the championship. Similar to Part I, we focus on how each prefetch algorithm functions, and why it is effective. Finer implementation details can be obtained [&#8230;]]]>
</description>
				<content:encoded><![CDATA[<div><img width="300" height="187" src="https://www.sigarch.org/wp-content/uploads/2026/04/DPC-4-Part-1-300x187.png" class="attachment-medium size-medium wp-post-image" alt="DPC-4 Concept Art (Indigo)" style="margin-bottom:15px;margin-left:15px;float:right;" decoding="async" loading="lazy" /></div><p>This article continues (and concludes) the discussion on the proceedings of <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://sites.google.com/view/dpc4-2026/home?pli=1">DPC-4</a>, covering the remaining four contestants and a summary of the trends observed in all eight prefetchers presented in the championship. Similar to <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://www.sigarch.org/fourth-data-prefetching-championship-part-i/">Part I</a>, we focus on how each prefetch algorithm functions, and why it is effective. Finer implementation details can be obtained from the workshop <a class="WKVSfLCavKyjywFEXwZDFLAfdDosdiAqrY " href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://github.com/CMU-SAFARI/DPC4/tree/main/final-versions" target="_self" data-test-app-aware-link="">papers</a> or the source <a class="WKVSfLCavKyjywFEXwZDFLAfdDosdiAqrY " href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://github.com/CMU-SAFARI/DPC4/tree/main/submissions" target="_self" data-test-app-aware-link="">code</a>.</p>
<h3><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://github.com/CMU-SAFARI/DPC4/blob/main/final-versions/BertiGO-final.pdf"><strong>BertiGO</strong> (</a><em>Simranjit Singh, University of Murcia; Agustín Navarro Torres, University of Zaragoza; Alberto Ros (University of Murcia</em></h3>
<h4 id="ember675" class="ember-view reader-text-block__heading-3">Motivation</h4>
<p id="ember676" class="ember-view reader-text-block__paragraph">When evaluating the baseline prefetcher configuration, the authors noted that <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://dpc3.compas.cs.stonybrook.edu/pdfs/Berti.pdf">Berti</a> frequently issues redundant prefetch requests for lines already prefetched or present in the cache. Also, using only the PC provides very limited context for pattern recognition, limiting the prediction capabilities of Berti. Furthermore, <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://dl.acm.org/doi/10.1145/3466752.3480114">Pythia</a> is found to generate a lot of useless prefetches for some workloads, which pollutes the L2 cache and wastes memory bandwidth.</p>
<h4 id="ember677" class="ember-view reader-text-block__heading-3">Idea</h4>
<ol>
<li>A Region-Based Bit-Map Filter is added, which is a fully associative structure storing the prefetched and accessed cache lines per region, in the form of a bit-vector. For regions tracked by the filter, having the M-th bit set in the bitmap implies that we drop all prefetch requests for the M-th cache line inside the region.</li>
<li>In addition to using PC, the authors propose using a hash (shifted XOR) of the last 4 PCs with the current PC, to index the Berti tables with additional context.</li>
<li>Set-Dueling is added to Pythia: instead of using the default policy to issue prefetches, 5 different policies are introduced, including a No-Prefetch policy that disables Pythia. All 5 policies are enabled for a 10M-instruction tournament, at the end of which the policy with the lowest miss rate is chosen for the rest of execution.</li>
<li>An Adaptive Next Line (ANeLin) is added to the LLC, which uses a sampling cache to track the demand misses and insert next-line prefetches. A heuristic mechanism is used to track useful and useless prefetches globally and per-PC. ANeLin can be disabled if the ratio of useful to useless prefetches drops below a threshold.</li>
</ol>
<h4 id="ember679" class="ember-view reader-text-block__heading-3">Why It Works</h4>
<p id="ember680" class="ember-view reader-text-block__paragraph">Adding a Bit-Map Filter eliminates redundant and useless prefetches. Using PC history adds context from the program flow while learning memory accesses with minimal overhead. Disabling Pythia and Next Line prefetching when they do not generate enough useful prefetches solves the problem of cache pollution due to wasteful prefetching. This is especially useful in the constrained bandwidth and multicore scenarios where data and memory need to be shared judiciously for optimal performance.</p>
<h3 id="ember682" class="ember-view reader-text-block__heading-2"><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://github.com/CMU-SAFARI/DPC4/blob/main/final-versions/EDP-final.pdf"><strong>Entangling Data Prefetcher</strong> (</a><em>Agustín Navarro Torres, Universidad de Zaragoza;  Simranjit Singh, University of Murcia; Biswabandan Panda, IIT Bombay; Alberto Ros, University of Murcia)</em></h3>
<h4 id="ember684" class="ember-view reader-text-block__heading-3">Motivation</h4>
<p id="ember685" class="ember-view reader-text-block__paragraph">Comparing Berti with other state-of-art prefetchers, the authors identify a SPEC2017 workload where Berti achieves negligible performance gain over no-prefetch baseline. Profiling this trace reveals that it consists of long-reuse strides (stride accesses separated by 2K-cycle interval) and zero-strides (consecutive accesses to the same cache line). Berti cannot issue zero-delta prefetches, and even though prefetches are correctly issued for long-reuse deltas, they get evicted before the cache line gets accessed. T-SKID, a Time Skipping Prefetcher is built on top of a standard PC-Stride prefetcher, but decouples the PC that triggers a prefetch (TriggerPC) from the PC that trains the predictor(TargetPC). This allows it to prefetch long-reuse and zero stride patterns. However, the underlying stride prefetcher limits its scope to constant stride instead of complex delta patterns predicted easily by Berti.</p>
<h4 id="ember686" class="ember-view reader-text-block__heading-3">Idea</h4>
<p id="ember687" class="ember-view reader-text-block__paragraph">EDP is proposed as a VA-based L1D prefetcher. It gets trained and triggered on cache misses or prefetch hits (cache hit on a prefetched line). For every TargetPC, it records the fill latency of the demand access or prefetch request. It then searches the global PC history for the most recent PC that was observed more than (current cycle &#8211; fill latency) cycles ago – this is the TriggerPC which could have triggered a timely prefetch for TargetPC. This ‘Entangling Pair’ of PCs is added to the Entangling Table, that stores the set of TargetPCs for a given TriggerPC. EDP also looks at the address history of each TargetPC to calculate the list of timely deltas (similar to Berti) and stores them with the current address in a Delta Table indexed by TargetPC. To issue prefetches, the TriggerPC is used to obtain one or more TargetPC, which are used to obtain address and deltas for timely prefetch. The prefetches calculated in this way are passed through a Bloom Filter to drop redundant requests, and then placed in a Proxy Prefetch Queue (PPQ) where the prefetch request waits till slots open up in the demand read queue. If there is no space in the latter, prefetch requests are not issued. Pythia is implemented at L2, with a throttling mechanism at LLC that tracks each core&#8217;s requests and sets the EDP aggressiveness.</p>
<h4 id="ember688" class="ember-view reader-text-block__heading-3">Why It Works</h4>
<p id="ember689" class="ember-view reader-text-block__paragraph">Using a different PC to trigger prefetches allows EDP to successfully prefetch zero and long reuse delta patterns for its target PC. Filtering out redundant prefetches reduces contention for resources. Using a dedicated PPQ for prefetch requests prevents prefetches from competing with critical loads for resources. The LLC throttling mechanism helps evenly distribute resources in the multi-core scenario.</p>
<h3 id="ember691" class="ember-view reader-text-block__heading-2"><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://github.com/CMU-SAFARI/DPC4/blob/main/final-versions/uMAMA-final.pdf"><strong>Composite Prefetching with Bandits</strong> (</a><em>Charles Block, Pedro Palacios, Abraham Farrell, Gerasimos Gerogiannis, Josep Torrellas, University of Illinois at Urbana-Champaign)</em></h3>
<h4 id="ember693" class="ember-view reader-text-block__heading-3">Motivation</h4>
<p id="ember694" class="ember-view reader-text-block__paragraph">The authors point out that the current state-of-the-art prefetchers try to optimize low-level metrics such as accuracy, timeliness and coverage. The system performance (IPC) depends on these factors, but can have variable sensitivity to each of them depending on the workload and program phase. Furthermore, a single prefetcher is generally insufficient to deliver the best performance for a diverse set of workloads – industrial processors generally deploy a composite prefetcher consisting of multiple prefetch engines.</p>
<h4 id="ember695" class="ember-view reader-text-block__heading-3">Idea</h4>
<p id="ember696" class="ember-view reader-text-block__paragraph">A Multi-Armed Bandit is a Reinforcement Learning agent that chooses the best action (arm) to maximize the reward function value. Inspired by this, a Micro-Armed Bandit (MAB) is used to prefetch at L2C. Each ‘arm’ consists of different configurations for 5 state-of-the-art prefetchers-</p>
<ul>
<li>Next Line, Spatial Memory Streaming, Best Offset Prefetcher: Can be turned ON or OFF</li>
<li>Stride, Stream prefetchers: Degree can be tuned to control aggressiveness</li>
</ul>
<p id="ember698" class="ember-view reader-text-block__paragraph">A bloom filter is implemented to prevent issuing redundant prefetches. Each arm is used for a fixed time period (bandit step) after which the reward generated by it is evaluated by the agent. This is evaluated against the rewards generated previously to calculate which arm to use next. The total IPC of the core is used as a reward function for the MAB.</p>
<p id="ember699" class="ember-view reader-text-block__paragraph">To optimize multi-core performance, another agent called ‘µMama’ is added at the system level, using the geometric mean of IPCs across all cores as a reward function. At each timestep, it decides whether to allow the cores to pursue their independent actions, or to force them into joint actions which have a record of increasing the µMama reward.</p>
<h4 id="ember700" class="ember-view reader-text-block__heading-3">Why It Works</h4>
<p id="ember701" class="ember-view reader-text-block__paragraph">Using Reinforcement Learning to directly maximize the system performance ensures that the prefetcher dynamically re-configures itself with execution to improve IPC. The caveat is that this now becomes a search space problem &#8211; the arms of the bandit need to be diverse enough to support different kinds of workloads, in order to deliver the best performance.</p>
<h3 id="ember703" class="ember-view reader-text-block__heading-2"><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://github.com/CMU-SAFARI/DPC4/blob/main/final-versions/GBerti-final.pdf"><strong>Global Berti</strong> (</a><em>Gilead Posluns, Mark Jeffrey; University of Toronto)</em></h3>
<h4 id="ember705" class="ember-view reader-text-block__heading-3">Motivation</h4>
<p id="ember706" class="ember-view reader-text-block__paragraph">Berti is a state-of-the-art prefetcher that detects Streaming patterns, i.e., consistent delta values between accesses by the <em>same</em> PC. Practical workloads however, often exhibit Spatial patterns identified by consistent delta values between accesses by <em>different </em>PCs. In the absence of streaming patterns, prefetching based on spatial patterns could alleviate the efficacy of Berti.</p>
<h4 id="ember707" class="ember-view reader-text-block__heading-3">Idea</h4>
<p id="ember708" class="ember-view reader-text-block__paragraph">Global Berti detects spatial patterns using Berti’s existing structures – the History Table conventionally stores within a row, the addresses of all the lines accessed by a particular PC, in FIFO order. When a streaming pattern cannot be detected, local training is useless and Global Berti looks at the most recent address for all PCs to detect spatial patterns (global training). Berti’s Delta Table holds the row delta values for the same PC; Global Berti stores the global deltas (across PCs) in the same table, adding a local bit to differentiate between streaming and spatial training.</p>
<h4 id="ember709" class="ember-view reader-text-block__heading-3">Why It Works</h4>
<p id="ember710" class="ember-view reader-text-block__paragraph">By itself, Berti is quite effective at detecting and covering streaming patterns. Adding the capability to detect spatial patterns in the absence of streaming patterns increases Global Berti’s coverage and therefore, the overall performance. As expected, the highest speedup over Berti is obtained in SPEC2017 and Graph workloads that are dominated by irregular accesses which require spatial prefetching. On the other hand, AI workloads containing mostly streaming patterns see a much lesser speedup.</p>
<h3 id="ember712" class="ember-view reader-text-block__heading-2">General Trends</h3>
<p id="ember713" class="ember-view reader-text-block__paragraph">Although the major focus of almost all DPC-4 submissions is to overcome the limitations of the high-performing Berti/Pythia baseline, they highlight several key trends in data prefetching research:</p>
<ul>
<li><strong>Prefetching across Physical Page Boundaries: </strong>Issuing page-crossing prefetches is extremely useful for AI workloads since they are dominated by streaming accesses. This is leveraged by most submissions to gain an edge over the baseline prefetcher configuration.</li>
<li><strong>Preventing Redundant Prefetches: </strong>Quite a few papers also combat excessive prefetching and resource contention through advanced throttling, priority, and filtering mechanisms.</li>
<li><strong>Increased System-Level and Multi-Core Awareness:</strong> There is a growing emphasis on system-aware solutions to judiciously manage shared resources like memory bandwidth, which is constrained in high-core-count datacenters. This includes core-level fairness throttling (Emender, EDP) and global coordination agents (µMama) to dynamically adjust prefetcher configurations for optimal multi-core performance.</li>
<li><strong>Expanding Pattern Coverage for Diverse Workloads:</strong> Submissions seek to improve coverage beyond simple streaming patterns. This includes detecting spatial patterns across different PCs (Global Berti), and targeting complex patterns like long-reuse and zero-strides (EDP). The adoption of PC history (BertiGO) also provides better context for pattern recognition.</li>
<li><strong>Shift Towards Adaptive and Composite Designs:</strong> Recognizing that a  single prefetcher is insufficient for diverse workloads, the trend moves toward composite prefetchers. This is accompanied by dynamic re-configuration to select the best prefetcher setting at runtime, and adaptive heuristics to tune aggressiveness.</li>
</ul>
<h3>About the Author</h3>
<p><span style="font-weight: 400;">Digvijay Singh obtained his Bachelor’s degree from BITS Pilani and his Master’s degree from Texas A&amp;M University where he worked on data prefetching as part of the CAMSIN research group. He currently works as a Silicon Architect in Google’s mobile CPU team.</span></p>
<p class="disclaim"><strong>Disclaimer:</strong> <em>These posts are written by individual contributors to share their thoughts on the Computer Architecture Today blog for the benefit of the community. Any views or opinions represented in this blog are personal, belong solely to the blog author and do not represent those of ACM SIGARCH or its parent organization, ACM.</em></p><Img align="left" border="0" height="1" width="1" alt="" style="border:0;float:left;margin:0;padding:0;width:1px!important;height:1px!important;" hspace="0" src="https://feeds.feedblitz.com/~/i/954807320/0/sigarch-cat">
]]>
</content:encoded>
			<wfw:commentRss>https://feeds.feedblitz.com/~/954807320/0/sigarch-cat/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
	<post-id xmlns="com-wordpress:feed-additions:1">103014</post-id></item>
<item>
<feedburner:origLink>https://www.sigarch.org/fourth-data-prefetching-championship-part-i/</feedburner:origLink>
		<title>Fourth Data Prefetching Championship: Part I</title>
		<link>https://feeds.feedblitz.com/~/954636863/0/sigarch-cat/</link>
		<comments>https://feeds.feedblitz.com/~/954636863/0/sigarch-cat/#respond</comments>
		<pubDate>Mon, 27 Apr 2026 14:00:53 +0000</pubDate>
		<dc:creator><![CDATA[Digvijay Singh]]></dc:creator>
		<category><![CDATA[ACM SIGARCH]]></category>
		<category><![CDATA[Data Prefetcher]]></category>
		<category><![CDATA[Memory Wall]]></category>
		<guid isPermaLink="false">https://www.sigarch.org/?p=103010</guid>
		<description><![CDATA[<div><img width="300" xheight="187" src="https://www.sigarch.org/wp-content/uploads/2026/04/DPC-4-Part-2-300x187.png" class="attachment-medium size-medium wp-post-image" alt="DPC-4 Concept Art (Blue)" style="max-width:100% !important;height:auto !important;margin-bottom:15px;margin-left:15px;float:right;"  loading="lazy" /></div>This article is the first in a two-part series that summarizes the key contributions of 4th Data Prefetching Championship (DPC-4), held in conjunction with the 32nd iteration of HPCA in 2026. While discussing innovative data prefetching techniques presented in this contest, we focus on the functionality of proposed algorithms and also explain why they are [&#8230;]]]>
</description>
				<content:encoded><![CDATA[<div><img width="300" height="187" src="https://www.sigarch.org/wp-content/uploads/2026/04/DPC-4-Part-2-300x187.png" class="attachment-medium size-medium wp-post-image" alt="DPC-4 Concept Art (Blue)" style="margin-bottom:15px;margin-left:15px;float:right;" decoding="async" loading="lazy" /></div><p class="ember-view reader-text-block__paragraph">This article is the first in a two-part series that summarizes the key contributions of 4th Data Prefetching Championship (<a class="WKVSfLCavKyjywFEXwZDFLAfdDosdiAqrY " href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://sites.google.com/corp/view/dpc4-2026/home" target="_self" data-test-app-aware-link="">DPC-4</a>), held in conjunction with the 32nd iteration of <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://2026.hpca-conf.org/track/hpca-2026-main-conference">HPCA</a> in 2026. While discussing innovative data prefetching techniques presented in this contest, we focus on the functionality of proposed algorithms and also explain why they are effective. Finer implementation details can be found from the <a class="WKVSfLCavKyjywFEXwZDFLAfdDosdiAqrY " href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://github.com/CMU-SAFARI/DPC4/tree/main/final-versions" target="_self" data-test-app-aware-link="">papers</a> or the source <a class="WKVSfLCavKyjywFEXwZDFLAfdDosdiAqrY " href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://github.com/CMU-SAFARI/DPC4/tree/main/submissions" target="_self" data-test-app-aware-link="">code</a>.</p>
<h3 id="ember612" class="ember-view reader-text-block__heading-3">Implementation Constraints</h3>
<p id="ember613" class="ember-view reader-text-block__paragraph">All prefetchers are evaluated against a baseline configuration that employs: <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://dpc3.compas.cs.stonybrook.edu/pdfs/Berti.pdf">Berti</a> prefetcher (<a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://dpc3.compas.cs.stonybrook.edu/">DPC3</a> winner) at L1D (Level-1 Data cache) and <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://dl.acm.org/doi/10.1145/3466752.3480114">Pythia</a> prefetcher at L2 (Level-2 cache). While there were no constraint on design complexity, upper limits were defined on the storage budget of the prefetchers to ensure the design was practically feasible for implementation. These limits were defined as follows: L1D Prefetcher: 32KB, L2 Prefetcher: 128KB, LLC (Last Level Cache) Prefetcher: 256KB.</p>
<h3 id="ember618" class="ember-view reader-text-block__heading-2">Keynotes</h3>
<p>The event included two keynote talks. The first keynote, titled &#8220;Is Prefetcher Research Still Alive?&#8221;, was given by <em><a id="ember621" class="ember-view" href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://www.linkedin.com/in/leeor-peled-125365b4/">Leeor Peled</a> </em>from Huawei. Leeor discussed the modern relevance of prefetching research, offering a pragmatic philosophy for academic researchers. He argued that the primary objective should not necessarily be to surpass &#8220;best-in-class&#8221; models – which are often the result of years of ‘engineered’ fine-tuning – but rather to introduce <strong>novel, high-potential concepts</strong> that invite further optimization. He emphasized that while an individual effort might not immediately surpass the state-of-the-art, a sufficiently &#8220;interesting&#8221; technique can evolve into a transformative solution through subsequent community-driven iteration.</p>
<p id="ember623" class="ember-view reader-text-block__paragraph">He suggested two optimizations that can be explored:</p>
<ol>
<li>Building a Semantic Prefetcher that correlates memory accesses with address generating code, i.e., a high-precision version of the Runahead Prefetcher that selectively runs only the code responsible for generating a future address.</li>
<li>Training neural networks to identify deep correlations between memory accesses, potentially unlocking the ability to predict complex, non-linear patterns that remain invisible to current heuristic-based logic.</li>
</ol>
<p id="ember625" class="ember-view reader-text-block__paragraph">The following issues can (and should) be addressed to build better prefetchers:</p>
<ul>
<li>Generalizing complex patterns, e.g. pointer chasing loads</li>
<li>Accurately choosing memory access with high correlation for better training</li>
<li>Prefetching to the appropriate cache level to optimize for timeliness</li>
<li>Throttling prefetches for fairness amongst multiple cores</li>
<li>Using LLMs to process memory traces instead of text sequences</li>
</ul>
<p id="ember631" class="ember-view reader-text-block__paragraph">The second keynote, titled &#8220;Data Prefetching: A Datacenter Perspective&#8221;, was given by <em><a id="ember630" class="ember-view" href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://www.linkedin.com/in/akanksha-j-a8336884/">Akanksha J.</a> from Google. </em> Addressing the memory bottleneck problem in modern datacenters (40% of the CPU cycles are spent idling for memory responses) Akanksha highlighted that cloud environments are characterized by massive multi-threading and incessant context switching. In these scenarios, a single thread may migrate across multiple cores, while each core rotates through a vast &#8220;plethora&#8221; of applications. The Google workloads utilized in DPC-4 are a better representation of this reality, and are primarily frontend-bound. Without a sophisticated instruction prefetcher to streamline code delivery, the underlying bottlenecks in data prefetching remain obscured and impossible to solve. She also analyzed structural failures of current prefetching solutions, identifying these primary aspects:</p>
<ol>
<li>Current design philosophy focuses on &#8220;tuning for the common case,&#8221; resulting in hard-coded heuristic values—such as fixed confidence thresholds and prefetch degrees—that are taped out into non-programmable silicon. While these &#8220;black boxes&#8221; are meticulously engineered to squeeze every drop of performance from SPEC workloads, they lack the flexibility required for the high heterogeneity of datacenter tasks. Consequently, these resource-hungry techniques often penalize cloud performance rather than enhancing it.</li>
<li>If we disable hardware prefetchers entirely and rely on software to insert prefetches, we miss out on critical opportunities to utilize valuable information about system states (coherence, timeliness, cache hits/misses) that improves prefetching. Akanksha proposed a shift towards <strong>&#8220;Software-Defined Prefetching,&#8221;</strong> a paradigm that transcends current <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://www.arm.com/glossary/isa">ISA</a> limitations. In this model, the software layer dynamically selects which code segments to target and determines the optimal hardware prefetcher to activate for peak accuracy. Simultaneously, the hardware leverages real-time system state data to maximize coverage.</li>
</ol>
<p id="ember633" class="ember-view reader-text-block__paragraph">Furthermore, Akanksha advocated for evaluating all prefetching techniques within constrained-bandwidth environments, arguing that such stress tests better reflect the realities of modern compute environments.</p>
<p>Now, on to prefetcher designs themselves.</p>
<h3 id="ember635" class="ember-view reader-text-block__heading-2"><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://github.com/CMU-SAFARI/DPC4/blob/main/final-versions/VIP-final.pdf"><strong>Virtual Inter-Page Prefetcher</strong></a> (<em>Ho Je Lee, Won Woo Ro; Yonsei University)</em></h3>
<h4 id="ember637" class="ember-view reader-text-block__heading-3">Motivation</h4>
<ul>
<li>Analyzing the baseline prefetecher configuration, the authors observed that the L2 Prefetcher (Pythia) is more effective than the L1 Prefetcher (Berti) in reducing Misses Per Kilo Instructions (MPKI) for the Last Level Cache (LLC).</li>
<li>Since Pythia operates in the Physical Address (PA) space, it is not feasible to let it issue prefetches across page boundaries, as incorrect physical page access poses a security risk.</li>
<li>A roofline study shows that there is significant performance to be gained when Pythia is allowed to issue page-cross prefetches in the PA space. This advantage amplifies when it is granted visibility of the Virtual Address (VA) space, preventing incorrect page accesses.</li>
</ul>
<h4 id="ember639" class="ember-view reader-text-block__heading-3">Idea</h4>
<p id="ember640" class="ember-view reader-text-block__paragraph">VIP is implemented at L1 level,  but issues prefetches to the L2. It gets trained on L1 Misses by reading the {<a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://www.geeksforgeeks.org/operating-systems/what-is-program-counter/">PC</a>, VA} information off the packets sent to L1 MSHR. These are written to the VIP Stride Table that calculates the observed stride for a particular PC and stores it. If a stride value is repeated, the confidence gets incremented. Otherwise it gets reset. The confidence value determines the prefetch degree.</p>
<h4 id="ember641" class="ember-view reader-text-block__heading-3">Why It Works</h4>
<p id="ember642" class="ember-view reader-text-block__paragraph">The implemented VIP configuration is a simple yet elegant solution to gain performance over the baseline by supplementing the existing Berti and Pythia prefetchers with cross-page prefetches (note that the DPC-3 version of Berti operates in the PA space and cannot issue prefetches across page boundaries). As expected, the stride prefetcher boosts AI workloads with sequential accesses of large data structures that span across pages. The typical CPU workloads such as SPEC see a moderate gain; the control-flow dominated Google workloads have a marginal slowdown since they rarely have uninterrupted streams.</p>
<h3 id="ember644" class="ember-view reader-text-block__heading-2"><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://github.com/CMU-SAFARI/DPC4/blob/main/final-versions/SPPAM-final.pdf"><strong>Signature Pattern Prediction and Access-Map Prefetcher</strong></a> (<em>Maccoy Merrell, Lei Wang, Paul Gratz, Stavros Kalafatis; Texas A&amp;M University)</em></h3>
<h4 id="ember646" class="ember-view reader-text-block__heading-3">Motivation</h4>
<p id="ember647" class="ember-view reader-text-block__paragraph">Access Map Pattern Matching (<a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://dl.acm.org/doi/10.1145/1542275.1542349">AMPM</a>) and Signature Path Prefetching (<a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://dl.acm.org/doi/10.5555/3195638.3195711">SPP</a>) are both considered state-of-the-art prefetching techniques; while SPP is sensitive to the order of memory accesses, AMPM is resistant to OoO execution. However, AMPM relies heavily on stored patterns for each region and  is unable to issue prefetches for new regions or when the observed accesses deviate from expectations. SPP excels at this and can even make predictions from its issued prefetches.</p>
<h4 id="ember648" class="ember-view reader-text-block__heading-3">Idea</h4>
<p id="ember649" class="ember-view reader-text-block__paragraph">Implemented at L2 level, a Region Table (RT) tracks all access maps (as bit-vectors) on a per-region basis. Upon a memory access, an N-bit portion from the respective access map is used to index a Pattern Table (PT). The PT outputs the most frequently occurring N-bit pattern as a prefetch candidate, which can be used to speculatively index the PT. Similar to SPP, speculative prefetching continues till the overall confidence drops below a threshold. The RT access map indicates the recently accessed cache lines and filters out redundant prefetches.</p>
<h4 id="ember650" class="ember-view reader-text-block__heading-3">Why It Works</h4>
<p id="ember651" class="ember-view reader-text-block__paragraph">The authors have identified the complementary nature of SPP and AMPM, and have combined them effectively to utilize the OoO resistance of AMPM with the Speculative mechanism of SPP. Additionally, numerous throttling mechanisms are implemented which consider pattern usefulness as well as global metrics such as DRAM bandwidth and overall usefulness to drop prefetches and set prefetch degree. SPPAM is implemented at L2C with Berti (the MICRO version which operates in the VA space) at L1D and Bingo at LLC. Similar to the previous paper, the cross-page stream information is passed to SPPAM from L1D.</p>
<h3 id="ember653" class="ember-view reader-text-block__heading-2"><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://github.com/CMU-SAFARI/DPC4/blob/main/final-versions/Emender-final.pdf"><strong>Emender</strong></a> (<em>Jiajie Chen, Tingji Zhang, Xiaoyi Liu, Xuefeng Zhang, Peng Qu, Youhui Zhang; Tsinghua University)</em></h3>
<h4 id="ember655" class="ember-view reader-text-block__heading-3">Motivation</h4>
<p id="ember656" class="ember-view reader-text-block__paragraph">An evaluation of different combinations of state-of-the-art prefetchers shows that <a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://dl.acm.org/doi/10.1109/MICRO56248.2022.00072">VBerti</a> (L1D) and Pythia (L2) is the highest performing combination. Here, VBerti refers to the MICRO version of Berti that operates in the VA space, allowing it to issue page-crossing prefetches. It is observed that this optimal prefetcher combination issues too many prefetch requests that fill the prefetch queue quickly, which leads to useful prefetches getting dropped. A second-order effect of a full prefetch queue is the excessive usage of L1D to Memory bandwidth that can delay critical loads.</p>
<h4 id="ember657" class="ember-view reader-text-block__heading-3">Idea</h4>
<p id="ember658" class="ember-view reader-text-block__paragraph">Four key features are added to tackle the problem of over-prefetching in the VBerti+Pythia configuration:</p>
<ol>
<li>Pending Target Buffer is added to sort all issued prefetches by confidence, which helps prioritize useful prefetches between different PCs.</li>
<li>Cuckoo Filter is added which tracks the VAs already present in the cache to prevent redundant prefetches. This structure is chosen due to its O(1) query time, high accuracy and zero false negatives.</li>
<li>Dynamic Confidence Threshold is added which increases with the cache miss rate, throttling low-confidence prefetches.</li>
<li>A Fairness-based Throttling scheme is implemented across cores, which tracks the useless prefetches per-core at L3 and stops the core with the most useless prefetches from prefetching.</li>
</ol>
<h4 id="ember661" class="ember-view reader-text-block__heading-3">Why It Works</h4>
<p id="ember662" class="ember-view reader-text-block__paragraph">The authors identify problematic areas in the baseline Berti+Pythia system and propose features to effectively address them. The best performance improvement comes from the Cuckoo Filter for single-core and Fairness Throttling for multi-core configuration. Since Emender provides the least gain for limited bandwidth configuration, it would be interesting to look at the accuracy data.</p>
<h3 id="ember664" class="ember-view reader-text-block__heading-2"><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://github.com/CMU-SAFARI/DPC4/blob/main/final-versions/sBerti-final.pdf"><strong>sBerti</strong></a> (<em>Jiapeng Zhou, Ben Chen, Kunlin Li, Yun Chen; HKUST, Guangzhou)</em></h3>
<h4 id="ember666" class="ember-view reader-text-block__heading-3">Motivation</h4>
<p id="ember667" class="ember-view reader-text-block__paragraph">When profiling the DPC4 workloads on the given baseline prefetcher configuration (Berti + Pythia), the authors observed a high L1D miss rate in the AI-ML and Google workloads. A deeper analysis of the traces indicated that most of these misses occurred when the access stream moved across the 4KB physical page boundary, which happens frequently in these workloads. The version of Berti used in the baseline does not issue prefetches across page boundaries, and thus, a stride prefetcher can help.</p>
<h4 id="ember668" class="ember-view reader-text-block__heading-3">Idea</h4>
<p id="ember669" class="ember-view reader-text-block__paragraph">A decoupled Smart Stride Prefetcher is added at L1D, which operates on the VA space and can  therefore track memory access streams across page boundaries. It is trained using a Smart Stride Table (SST), which is indexed by a hash of the PC, and subtracts the lastVA from the current VA to calculate the delta value. If the absolute value of delta is a multiple of the stored stride, the confidence is updated; this also provides resistance to out-of-order execution. Prefetches are issued if this confidence is greater than a static threshold. The lookahead is tuned via a heuristic which is incremented upon observing late prefetches and decremented by timely prefetches. A Recent Prefetch Table stores the recently issued prefetches to track their timeliness and filter duplicate prefetches between Berti and Smart Stride engines.</p>
<h4 id="ember670" class="ember-view reader-text-block__heading-3">Why It Works</h4>
<p id="ember671" class="ember-view reader-text-block__paragraph">The addition of a decoupled stride prefetcher gives sBerti the ability to issue prefetches across physical page boundaries, reducing the “Cold-start Penalty” of Berti. The heuristic based dynamic distance adjustment helps tune the aggressiveness at runtime, allowing longer lookahead for AI-ML workloads dominated by streaming accesses. The final sBerti configuration (Stride + Berti at L1D, Pythia at L2) delivers the best performance in a full bandwidth scenario, where the stride engine can prefetch further ahead.</p>
<p>We will overview the rest of the prefetchers in part 2 of this post.</p>
<h3>About the Author</h3>
<p><span style="font-weight: 400;">Digvijay Singh received his Bachelor’s degree from BITS Pilani and his Master’s degree from Texas A&amp;M University where he worked on data prefetching as part of the CAMSIN research group. He currently works as a Silicon Architect in Google’s mobile CPU team.</span></p>
<p class="disclaim"><strong>Disclaimer:</strong> <em>These posts are written by individual contributors to share their thoughts on the Computer Architecture Today blog for the benefit of the community. Any views or opinions represented in this blog are personal, belong solely to the blog author and do not represent those of ACM SIGARCH or its parent organization, ACM.</em></p><Img align="left" border="0" height="1" width="1" alt="" style="border:0;float:left;margin:0;padding:0;width:1px!important;height:1px!important;" hspace="0" src="https://feeds.feedblitz.com/~/i/954636863/0/sigarch-cat">
]]>
</content:encoded>
			<wfw:commentRss>https://feeds.feedblitz.com/~/954636863/0/sigarch-cat/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
	<post-id xmlns="com-wordpress:feed-additions:1">103010</post-id></item>
<item>
<feedburner:origLink>https://www.sigarch.org/beyond-qubits-a-systems-view-of-hybrid-cv-dv-quantum-computing/</feedburner:origLink>
		<title>Beyond Qubits: A Systems View of Hybrid CV-DV Quantum Computing</title>
		<link>https://feeds.feedblitz.com/~/954105707/0/sigarch-cat/</link>
		<comments>https://feeds.feedblitz.com/~/954105707/0/sigarch-cat/#respond</comments>
		<pubDate>Mon, 20 Apr 2026 15:31:53 +0000</pubDate>
		<dc:creator><![CDATA[Yuan Liu, Zihan Chen, Shubdeep Mohapatra, Jim Furches, Zheng (Eddy) Zhang, Huiyang Zhou]]></dc:creator>
		<category><![CDATA[ACM SIGARCH]]></category>
		<category><![CDATA[Quantum Computing]]></category>
		<guid isPermaLink="false">https://www.sigarch.org/?p=102865</guid>
		<description><![CDATA[<div><img width="300" xheight="167" src="https://www.sigarch.org/wp-content/uploads/2026/04/Picture1-300x167.png" class="attachment-medium size-medium wp-post-image" alt="" style="max-width:100% !important;height:auto !important;margin-bottom:15px;margin-left:15px;float:right;"  loading="lazy" /></div>Hybrid continuous-discrete-variable (CV-DV) quantum computing combines oscillators and qubits to tackle problems that are difficult for either model alone, from bosonic simulation to quantum error correction. At ASPLOS 2026, our tutorial introduced the foundations, compilation stack, benchmarking methods, and programming tools behind this emerging architecture model. In this blog post, we overview the key elements [&#8230;]]]>
</description>
				<content:encoded><![CDATA[<div><img width="300" height="167" src="https://www.sigarch.org/wp-content/uploads/2026/04/Picture1-300x167.png" class="attachment-medium size-medium wp-post-image" alt="" style="margin-bottom:15px;margin-left:15px;float:right;" decoding="async" loading="lazy" /></div><p><b>Hybrid continuous-discrete-variable (CV-DV) quantum computing combines oscillators and qubits to tackle problems that are difficult for either model alone, from bosonic simulation to quantum error correction. At ASPLOS 2026, our tutorial introduced the foundations, compilation stack, benchmarking methods, and programming tools behind this emerging architecture model. In this blog post, we overview the key elements of our tutorial. </b></p>
<p><span style="font-weight: 400;">Tutorial website: </span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://cvdv.ncsu.edu/resources/asplos-tutorial/"><span style="font-weight: 400;">https://cvdv.ncsu.edu/resources/asplos-tutorial/</span></a></p>
<h3><b>Foundations</b></h3>
<p><span style="font-weight: 400;">We began with the foundations of hybrid CV-DV quantum computing, introducing the physical model, mathematical language, and programming abstractions behind qubit-oscillator systems. Many leading quantum platforms naturally combine qubits with oscillator modes, such as cavities, vibrational modes, or photonic fields. Rather than treating oscillators as auxiliary hardware, hybrid CV-DV computing views their large Hilbert spaces as a computational resource.</span></p>
<p><span style="font-weight: 400;">The tutorial covered core representations of CV states in both Fock space and phase space, along with the key operators and gate families that support universal CV-DV computation. A central message was that hybrid systems are not simply “qubits plus extra hardware,” but a distinct computational model with their own instruction sets, abstractions, and compilation challenges. We showed how familiar qubit concepts such as Pauli and Clifford structure extend into the oscillator setting through displacement operations, squeezing, quadratic Hamiltonians, beamsplitters, and controlled hybrid interactions.</span></p>
<p><span style="font-weight: 400;">We also discussed why this matters from a computer architecture perspective. Hybrid CV-DV systems introduce new instruction set architectures (ISAs), abstract machine models (AMMs), and compilation choices that help separate hardware details from software design. Depending on the platform and compiler stack, the same computation may be expressed in phase-space language, Fock-space language, or a mixed qubit-oscillator representation.</span></p>
<p><span style="font-weight: 400;">To ground these ideas, we highlighted emerging algorithmic primitives and applications where hybrid systems may offer advantages, including oscillator-mediated entangling gates, state-transfer protocols, Hamiltonian simulation, bosonic quantum error correction, vibronic dynamics, and sensing. We closed the session by surveying two leading implementation pathways, superconducting circuit QED and trapped-ion systems, and discussing the distinct control and connectivity tradeoffs they expose. A comprehensive tutorial on the foundations of hybrid CV-DV quantum processors is available </span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://journals.aps.org/prxquantum/abstract/10.1103/4rf7-9tfx"><span style="font-weight: 400;">here</span></a><span style="font-weight: 400;">.</span></p>
<h3><b>Compilation</b></h3>
<p><span style="font-weight: 400;">We also presented Strategies and Tools to Compile CV-DV Quantum Circuits. We began by emphasizing why Hamiltonian simulation is a central application and one of the most promising directions for hybrid continuous-variable and discrete-variable (CV-DV) quantum systems. CV systems can naturally represent continuous degrees of freedom, while DV systems provide strong control and interaction structures. Together, they enable important applications in areas such as quantum chemistry and materials science. However, a key challenge lies in decomposing the time-evolution operator e^{-iHt} into a sequence of executable quantum gates. This transformation is fundamentally a compilation problem, bridging high-level quantum algorithms and low-level hardware. As such, compilers play a critical role in hybrid quantum systems.</span></p>
<p><span style="font-weight: 400;">We then focused on the dominant approach today: symbolic compilation. In particular, we discussed two early CV-DV Hamiltonian simulation compilers from </span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://dl.acm.org/doi/full/10.1145/3695053.3731065"><span style="font-weight: 400;">Chen et al., ISCA’25</span></a><span style="font-weight: 400;"> and </span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://ieeexplore.ieee.org/abstract/document/11250346"><span style="font-weight: 400;">Decker et al., QCE’25</span></a><span style="font-weight: 400;">. The core idea is to avoid direct matrix-based computation and instead leverage the algebraic structure of operators for rule-based decomposition. Techniques such as Trotter-Suzuki product formulas, the Baker–Campbell–Hausdorff (BCH) expansion, and bosonic commutation relations are used to gradually break down complex Hamiltonians into hardware-executable primitive gates. This process is typically implemented through rule matching and recursive rewriting, where expressions are repeatedly transformed until only supported base gates remain. While this approach avoids the exponential blowup of high-dimensional matrices, it introduces tradeoffs between approximation error and resource overhead.</span></p>
<p><span style="font-weight: 400;">Finally, we analyzed the limitations of current compilers and outlined future research directions. Key challenges include limited gate sets and decomposition rules, the tradeoff between accuracy and resource cost, hardware connectivity constraints, and insufficient optimization flexibility. To address these issues, we highlighted the need for improved programmability, richer native gate support, more accurate cost models, and optimizations that exploit algebraic properties such as commutativity. We also presented the </span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://github.com/ruadapt/Genesis-CVDV-Compiler"><span style="font-weight: 400;">Genesis compiler</span></a><span style="font-weight: 400;"> from Chen et al., ISCA’25 as an end-to-end solution example, including typical use cases and code snippets. Genesis employs a multi-level intermediate representation (IR) and a full compilation pipeline to automatically translate Hamiltonians into limited hardware connectivity physical circuits, demonstrating a systematic and extensible compilation framework for hybrid CV-DV quantum computing.</span></p>
<h3><b>Benchmark and Circuit Simulator</b></h3>
<p><span style="font-weight: 400;">We also presented </span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://arxiv.org/abs/2603.04398"><b>HyQBench</b></a> <span style="font-weight: 400;">by Mohapatra et al., an open-source benchmark suite implemented in Bosonic Qiskit and QuTiP. HyQBench covers eight representative hybrid circuits spanning three abstraction levels: primitives, algorithms, and applications. These include cat state generation, GKP state preparation, CV-to-DV state transfer, </span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://ieeexplore.ieee.org/abstract/document/11129874"><span style="font-weight: 400;">CV-DV QFT</span></a><span style="font-weight: 400;">, </span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://arxiv.org/abs/2501.11735"><span style="font-weight: 400;">CV-DV VQE</span></a><span style="font-weight: 400;">, CV-QAOA, Jaynes-Cummings-Hubbard (JCH) Hamiltonian simulation, and </span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://www.nature.com/articles/s41467-025-67694-5"><span style="font-weight: 400;">Shor’s algorithm</span></a><span style="font-weight: 400;">.</span></p>
<p><span style="font-weight: 400;">One key takeaway is that hybrid architectures can reduce hardware resources dramatically for some workloads. For example, simulating a 3-site JCH model in a DV-only encoding requires 9 qubits and 393 CNOT gates, whereas a hybrid implementation uses only 3 qumodes, 3 qubits, and 8 gates. This kind of reduction highlights why benchmarking hybrid systems requires more than simply counting qubits.</span></p>
<p><span style="font-weight: 400;">To support this, we introduced a feature map tailored to hybrid systems. In addition to standard structural metrics such as gate counts, circuit depth, and qubit/qumode counts, we proposed three CV-DV-specific metrics: Wigner negativity as a proxy for non-classicality and classical simulation hardness, truncation cost to quantify population near the Fock cutoff, and maximum energy. These metrics help separate workloads with very different simulation and execution behavior. For example, JCH simulation remains relatively close to Gaussian behavior, while CV-QAOA and Shor’s algorithm exhibit higher Wigner negativity and are harder to simulate classically.</span></p>
<p><span style="font-weight: 400;">We also discussed early hardware validation. A cat-state preparation benchmark was executed on Sandia National Laboratories’ QSCOUT trapped-ion platform and achieved a fidelity of 0.71. HyQBench was further used to calibrate conditional displacement gates on the same platform, reinforcing the need for standardized benchmark suites that support both evaluation and device calibration. The full paper is available at </span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://arxiv.org/abs/2603.04398"><span style="font-weight: 400;">https://arxiv.org/abs/2603.04398</span></a><span style="font-weight: 400;">.</span></p>
<p><span style="font-weight: 400;">To lower the barrier to entry for this area, we also developed </span><b>HyQSim</b><span style="font-weight: 400;">, a browser-based hybrid CV-DV circuit simulator that requires no installation. HyQSim supports drag-and-drop circuit construction, arbitrary Fock cutoffs, and built-in visualization through Wigner plots, Fock-state amplitudes, and Bloch sphere views. It is available at </span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://cvdv.ncsu.edu/resources/simulator/"><span style="font-weight: 400;">https://cvdv.ncsu.edu/resources/simulator/</span></a><span style="font-weight: 400;">, and the code is hosted at </span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://github.com/shubdeepmohapatra01/HyQSim/"><span style="font-weight: 400;">https://github.com/shubdeepmohapatra01/HyQSim/</span></a><span style="font-weight: 400;">.</span></p>
<h3><b>Programming</b></h3>
<p><span style="font-weight: 400;">Finally, we discussed programming support for hybrid CV-DV systems. Quantum programming languages and frameworks have developed many important ideas over the years, including </span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://dl.acm.org/doi/10.1007/11417170_26"><span style="font-weight: 400;">linear quantum types</span></a><span style="font-weight: 400;"> for enforcing the no-cloning theorem, </span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://dl.acm.org/doi/10.1145/3385412.3386007"><span style="font-weight: 400;">automatic uncomputation</span></a><span style="font-weight: 400;"> of ancilla qubits, and </span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://dl.acm.org/doi/10.1145/2499370.2462177"><span style="font-weight: 400;">dynamic lifting of classical variables</span></a><span style="font-weight: 400;"> for mid-circuit measurement. Hybrid quantum computing introduces an additional requirement: heterogeneous quantum registers containing both qubits and qumodes.</span></p>
<p><span style="font-weight: 400;">To address this challenge, we developed </span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://arxiv.org/abs/2603.10919"><b>Hybridlane</b></a><span style="font-weight: 400;">, a CV-DV quantum programming framework built on </span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://arxiv.org/abs/1811.04968"><span style="font-weight: 400;">PennyLane</span></a><span style="font-weight: 400;">. By extending PennyLane, Hybridlane inherits a broad library of qubit algorithms, gates, and compilation routines while remaining familiar to existing users. Hybridlane tracks wire types automatically through symbolic circuit analysis and type inference, enabling scalable circuit construction, platform independence, and integration with downstream compilation flows.</span></p>
<p><span style="font-weight: 400;">The tutorial concluded with example workflows using Hybridlane. In one example, we reused an existing PennyLane quantum phase estimation template for a CV-DV Hamiltonian simulation and then lowered it through symbolic compilation to a gate sequence executable on the </span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://arxiv.org/abs/2209.11153"><span style="font-weight: 400;">Bosonic Qiskit</span></a><span style="font-weight: 400;"> backend. In another, we demonstrated a cross-platform workflow in which a conditional displacement gate was calibrated in simulation and then compiled for execution on Sandia’s QSCOUT trapped-ion platform. Together, these examples showed how hybrid quantum software can begin to support the same define-simulate-execute workflow that has become standard in mature qubit SDKs.</span></p>
<p><span style="font-weight: 400;">We hope Hybridlane helps enable a broader ecosystem of reusable software and research for hybrid quantum computing. It is available at </span><a href="https://feeds.feedblitz.com/~/t/0/0/sigarch-cat/~https://github.com/pnnl/hybridlane"><span style="font-weight: 400;">https://github.com/pnnl/hybridlane</span></a><span style="font-weight: 400;">.</span></p>
<h3><b>Closing</b></h3>
<p><span style="font-weight: 400;">Hybrid CV-DV computing sits at the intersection of quantum hardware, computer architecture, compilation, and programming systems. We hope this tutorial helps make the area more accessible to researchers across architecture, systems, programming languages, and quantum information, and we invite readers to explore the tutorial materials, benchmarks, and tools linked above.</span></p>
<h3><strong>About the Authors</strong></h3>
<p><span style="font-weight: 400;"><strong>Yuan Liu</strong> is an Assistant Professor of Electrical &amp; Computer Engineering and Computer Science at North Carolina State University. Prior to joining the NC State faculty, he was a postdoctoral researcher at the Massachusetts Institute of Technology. His research interests lie at the intersection of quantum computing, quantum engineering, quantum algorithms/architectures and applications.</span></p>
<p><span style="font-weight: 400;"><strong>Zihan Chen</strong> is a Ph.D. student in computer systems at Rutgers University, advised by Prof. Eddy Z. Zhang. His research focuses on compiler and system-level techniques, as well as parallel computing, to enhance the efficiency, programmability, scalability, and fault tolerance of emerging quantum computing systems.</span></p>
<p><span style="font-weight: 400;"><strong>Shubdeep Mohapatra</strong> is a Ph.D. candidate in Computer Engineering at NC State University, advised by Prof. Huiyang Zhou and Prof. Yuan Liu. His research focuses on quantum error characterization, mitigation, and benchmarking, aimed at improving the reliability and fault tolerance of near-term quantum computing systems.</span></p>
<p><span style="font-weight: 400;"><strong>Jim Furches</strong> is a post-masters research associate at Pacific Northwest National Laboratory. His current research interests are in quantum benchmarking, algorithms, and quantum programming and compilation.</span></p>
<p><span style="font-weight: 400;"><strong>Zheng (Eddy) Zhang</strong> is a Professor in the Department of Computer Science at Rutgers University. Her research focuses on full-stack compiler and programming systems for quantum computing. She studies how to better coordinate quantum applications, programming languages, intermediate representations, compilation, pulse-level control, and hardware architecture to improve the performance, usability, and scalability of quantum systems.</span></p>
<p><span style="font-weight: 400;"><strong>Huiyang Zhou</strong> is a Professor of Electrical and Computer Engineering at North Carolina State University. His current research interests include GPU architecture, processor security, and quantum computing.</span></p>
<p class="disclaim"><strong>Disclaimer:</strong> <em>These posts are written by individual contributors to share their thoughts on the Computer Architecture Today blog for the benefit of the community. Any views or opinions represented in this blog are personal, belong solely to the blog author and do not represent those of ACM SIGARCH or its parent organization, ACM.</em></p><Img align="left" border="0" height="1" width="1" alt="" style="border:0;float:left;margin:0;padding:0;width:1px!important;height:1px!important;" hspace="0" src="https://feeds.feedblitz.com/~/i/954105707/0/sigarch-cat">
]]>
</content:encoded>
			<wfw:commentRss>https://feeds.feedblitz.com/~/954105707/0/sigarch-cat/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
	<post-id xmlns="com-wordpress:feed-additions:1">102865</post-id></item>
</channel></rss>

