SUPERCOMPUTING NEWS SUPERCOMPUTING NEWS
    • MEDIA KIT
    • MOST READ
    • RSS FEED
    • ACADEMIA
    • AEROSPACE
    • APPLICATIONS
    • ASTRONOMY
    • AUTOMOTIVE
    • BIG DATA
    • BIOLOGY
    • CHEMISTRY
    • CLIENTS
    • CLOUD
    • DEFENSE
    • DEVELOPER TOOLS
    • EARTH SCIENCES
    • ECONOMICS
    • ENGINEERING
    • ENTERTAINMENT
    • GAMING
    • GOVERNMENT
    • HEALTH
    • OIL & GAS
    • INDUSTRY
    • INTERCONNECTS
    • MANUFACTURING
    • MIDDLEWARE
    • MOVIES
    • NETWORKS
    • PHYSICS
    • PROCESSORS
    • RETAIL
    • SCIENCE
    • STORAGE
    • SYSTEMS
    • VISUALIZATION
    • AcyMailing subscription form

    • ADD YOUR VIDEOS
    • MANAGE VIDEOS
    • CONVERSATION INBOX
    • SOCIAL ADVERTISER
    • SOCIAL NETWORK VIDEOS
    • SURVEYS
    • GROUPS
    • PAGES
    • MARKETPLACE LISTINGS
    • APPLICATIONS BROWSER
    • PRIVACY CONFIRM REQUEST
    • PRIVACY CREATE REQUEST
    • LEADERBOARD
    • POINTS LISTING
      • BADGES
    • TRADE SHOWS
Sign In
15.6 Microseconds, 156 simulations: Supercomputing maps the moving machinery of an enzyme
15.6 Microseconds, 156 simulations: Supercomputing maps the moving machinery of an enzyme
AI agents search 1.9 billion protein clusters, discover a new biological system
AI agents search 1.9 billion protein clusters, discover a new biological system
AI’s trillion dollar compute race hits a hard limit: There isn’t enough power
AI’s trillion dollar compute race hits a hard limit: There isn’t enough power
Supercomputing turns dark-matter waves into a testable prediction
Supercomputing turns dark-matter waves into a testable prediction
Alibaba’s superintelligence ambition puts supercomputing at the center of the AI race
Alibaba’s superintelligence ambition puts supercomputing at the center of the AI race
Japanese supercomputer recreates the birth of the Universe’s monster black holes
Japanese supercomputer recreates the birth of the Universe’s monster black holes
previous arrow
previous arrow
next arrow
next arrow
 
Shadow
15.6 Microseconds, 156 simulations: Supercomputing maps the moving machinery of an enzyme
Featured

15.6 Microseconds, 156 simulations: Supercomputing maps the moving machinery of an enzyme

CHRIS O'NEAL, PUBLISHER September 29, 2026, 8:00 am

For decades, structural biology has provided scientists with high-resolution snapshots of enzymes, characterizing molecular structures frozen in specific conformations via techniques such as X-ray crystallography. However, enzymes are dynamic systems that undergo continuous conformational changes, including the opening and closing of binding pockets, side-chain rotations, and the rearrangement of water molecules and substrates. Static crystal structures inherently fail to capture these transient states; therefore, extensive computational time is required to elucidate such motions.

A recent study published in ACS Omega illustrates the significant scale of molecular-dynamics (MD) computation necessary to transform static structural data into a comprehensive portrait of enzyme behavior. Researchers investigating pyrimidine-nucleoside phosphorylase (PyNP) from Bacillus subtilis conducted an experiment involving 156 independent production trajectories, totaling 15.6 microseconds of MD simulation. 

Leveraging computational resources from the Research Center for Computational Science (RCCS) in Okazaki, Japan, alongside cloud GPU infrastructure from vast.ai, the team utilized GROMACS to analyze 13 ligands across four distinct structural states of the enzyme. Each ligand structure combination was subjected to three independent 100-nanosecond replicas, providing a robust statistical ensemble. This work underscores the transition of high-performance computing (HPC) from a mere accelerator of calculations to a critical tool for statistical sampling, ultimately allowing researchers to address a fundamental question in the field: how do enzymes function when allowed to exhibit their inherent molecular mobility?

From Four Structures to 156 Simulations

The computational campaign began with four representations of PyNP.

Three came from experimental structures, while a fourth was generated by energy minimization. They represented different positions along the enzyme’s conformational range:

  • 1BRW — closed
  • 5EP8min — closed-like
  • 5EP8orig — semi-open
  • 5OLN — open

The distinction is important because PyNP contains a flexible gate region involving residues 153–170. At its center is Tyr165, a residue positioned to move over the substrate-binding pocket as the enzyme changes between open and closed configurations.

The researchers then introduced 13 ligands into these four structural states.

That creates 52 ligand–structure combinations.

Each combination was simulated three times using independently randomized initial velocities.

The arithmetic is simple:

13 ligands × 4 conformations × 3 replicas = 156 production trajectories.

Every trajectory was 100 nanoseconds long.

Together, that produced:

156 × 100 ns = 15.6 microseconds of molecular-dynamics simulation.

Trajectories were saved every 50 picoseconds, producing 2,001 frames for each 100-nanosecond trajectory.

This is precisely the type of workload for which HPC infrastructure becomes valuable.

A single molecular-dynamics trajectory can show what happens to one molecular system under one set of initial conditions. A large ensemble allows researchers to ask whether an observed behavior persists across different molecules, conformations, and independent simulations.

The computational experiment therefore was not simply:

Run a simulation.

It was:

Run enough simulations to determine which behaviors survive statistical variation.

The Supercomputer Behind the Molecular Experiment

The production calculations used GROMACS 2025.2, a molecular-dynamics package designed for highly parallel computation.

The researchers used the AMBER ff99SB-ILDN force field for the protein, TIP3P water, and GAFF2 parameters for the ligands. The systems were solvated, neutralized, and equilibrated before production calculations were performed under NPT conditions at 300 K.

The production timestep was 2 femtoseconds, with Particle Mesh Ewald electrostatics and hydrogen-bond constraints.

At that timestep, a 100-nanosecond trajectory represents approximately 50 million integration steps.

Across 156 production trajectories, that corresponds to roughly 7.8 billion molecular-dynamics integration steps for the primary simulation campaign.

That number is not itself a measure of scientific value, but it illustrates the computational scale behind the experiment.

The researchers generated the simulations primarily on the RCCS supercomputer. For the 5EP8orig structural state, one replica was generated on RCCS while two were generated using vast.ai cloud GPUs, providing both additional computational capacity and an opportunity to examine consistency across platforms.

The study therefore represents a hybrid HPC workflow: dedicated research-supercomputing resources supplemented by cloud GPU computation.

The computational output was substantial enough that the authors deposited all 156 raw trajectories in Zenodo, along with analysis scripts.

Why 156 Trajectories Matter

The key HPC insight is that molecular simulation has a sampling problem.

An enzyme’s behavior cannot necessarily be inferred from a single trajectory. Molecular dynamics is deterministic once its initial conditions are defined, but different starting velocities can produce different microscopic histories.

That is why the study used three independent replicas for every ligand structure combination.

The researchers then analyzed the trajectories using several metrics, including:

  • ligand-to-active-site distances;
  • the fraction of frames in which ligands remained associated with the active site;
  • residue-by-residue contact frequencies;
  • MM-PB(GB)SA binding-energy estimates; and
  • classifications describing ribose versus 2′-deoxyribose preferences.

This is where the HPC workload becomes scientifically meaningful.

The computer is not merely producing molecular movies.

It is producing a large statistical ensemble from which the researchers can extract patterns.

The Active Site Is Not Static

The computational results showed that the active-site pocket changes substantially between conformational states.

The calculated pocket volumes ranged from approximately 1,424 ų for 5EP8min to 2,549 ų for the open 5OLN structure.

The fully open state therefore has a pocket approximately 1.4 times larger than the closed 1BRW reference.

The authors interpret the structural progression as a transition from a contracted, substrate-trapping configuration toward an expanded, substrate-accessible configuration.

That observation establishes the computational problem.

If the pocket itself is changing size and shape, then ligand behavior cannot be understood simply by examining where a molecule sits in a single crystal structure.

The molecular machine has to be watched while it moves.

Tyr165 Emerges as a Molecular Gate

The most striking computational signal involved Tyr165.

Across the trajectory ensemble, Tyr165 showed a contact fraction of approximately 42.5% in the closed group, compared with 21.2% in the open group.

That represents a difference of 21.3 percentage points, with the reported statistical comparison producing a p-value of 1.2 × 10⁻⁵.

The physical interpretation is intuitive.

In the closed configuration, Tyr165 can sit over the active-site region, behaving like a molecular lid.

As the enzyme opens, that residue moves away from the ligand-binding region and toward solvent exposure.

The supercomputing campaign therefore converts a structural hypothesis into a trajectory-level observation:

The enzyme’s gate residue is dynamically coupled to its conformational state.

But the paper makes an important qualification.

The strongest Tyr165 signal comes primarily from 1BRW, which is a closed-state structure from the related species Geobacillus stearothermophilus, rather than from a closed-state crystal structure of the B. subtilis enzyme.

When 1BRW is excluded, the effect becomes substantially weaker.

The authors therefore do not present the result as definitive proof that Tyr165 universally behaves as a closed-state lid in B. subtilis. Instead, they identify it as a computationally supported mechanism that needs experimental confirmation using a closed-state structure from the same species.

That restraint is important.

The HPC calculation reveals a compelling molecular pattern, but the quality of the conclusion depends on the quality and comparability of the structures being sampled.

When Supercomputing Becomes a Hypothesis Generator

One of the most interesting aspects of the work is what happened after the initial 156 simulations.

The researchers used additional simulations to test whether the computationally identified mechanism could be challenged.

They performed three classes of additional all-atom molecular dynamics:

Tyr165 and other alanine mutants.

The researchers removed selected residues computationally to test predicted effects on ligand retention.

For the Y165A mutant, replacing Tyr165 with alanine substantially reduced ligand retention. The reported median late-window binding fraction was 0.02, compared with means of 0.50 for the Q153A control mutant and 0.37 for the K81A/K108A/K188A triple mutant. The reported one-sided Mann–Whitney comparison gave p = 0.049.

That is significant not because a computer has proven the biological mechanism, but because the simulation produced a falsifiable prediction.

The researchers explicitly characterize these simulations as computational support rather than experimental validation.

Adding Phosphate Changes the Question

The study also demonstrates an important limitation of classical molecular dynamics.

The initial simulations focused on enzyme–ligand complexes without the cosubstrate phosphate.

The researchers subsequently introduced an HPO₄²⁻ ion near the phosphate-binding region formed by Lys81, Lys108 and Lys188.

The phosphate interacted strongly with that lysine cluster and could approach the substrate’s anomeric carbon.

But the geometry required for the actual chemical substitution reaction was rarely observed.

The reactive in-line geometry occurred in less than 0.5% of contact frames, and in 30 of 33 systems it did not occur at all.

This result highlights a fundamental boundary between molecular dynamics and quantum chemistry.

Classical MD can model the movement and interactions of atoms using a predefined force field.

It cannot directly describe the breaking and making of chemical bonds at the electronic level.

The authors therefore point toward QM/MM and transition-state-level calculations as the next computational step.

For HPC researchers, that is an important distinction.

The computational problem does not end when the molecular dynamics run finishes.

Instead, one computational regime can identify the configurations that deserve to be examined by a more expensive method.

That is the essence of a hierarchical HPC workflow.

More Simulation Time Does Not Automatically Mean More Certainty

The study also extended selected central systems from 100 ns to 300 ns.

Those longer simulations produced retention and sugar-preference conclusions consistent with the original 100-nanosecond analysis.

But the authors appropriately qualify the result.

Only a subset of the systems was extended, and only about 46% of the original systems reached a plateau in the relevant analysis.

Consequently, the authors interpret the 100-nanosecond measurements primarily as measures of dynamic retention, rather than equilibrium affinity.

This is another important HPC lesson.

More compute is not automatically equivalent to complete sampling.

Molecular processes can occur on timescales much longer than those accessible to straightforward simulations.

Increasing a simulation from 100 to 300 nanoseconds may strengthen confidence in some observations without demonstrating that the system has reached thermodynamic equilibrium.

For the researchers, that distinction determines what the computation can legitimately claim.

From Molecular Movies to Mechanistic Maps

The study ultimately produces a more complicated picture than a simple “open versus closed” enzyme.

Different ligands interact with different combinations of active-site residues.

Gln153 emerges as an important anchor in some ligand pairs, while Lys108, Lys188, Gln153, and His82 form a broader interaction network in another case.

The computational data suggest that ribose-versus-2′-deoxyribose preferences can emerge from different residue combinations depending on the ligand and conformational state.

In other words, there is no single universal molecular switch controlling every ligand.

There is a network.

And uncovering that network is precisely where large-scale molecular simulation becomes useful.

A crystal structure can tell researchers where atoms are.

An ensemble of trajectories can begin to show how those atoms cooperate.

The HPC Pipeline Is the Experiment

Perhaps the most important lesson from the study for the supercomputing community is that the computer is not simply a supporting instrument.

The computational workflow is part of the experiment itself.

The researchers began with four molecular conformations.

They introduced 13 ligands.

They generated three independent trajectories for each combination.

They produced 156 simulations totaling 15.6 microseconds.

They analyzed thousands of molecular snapshots from those trajectories.

They identified a candidate molecular gate.

They computationally removed the gate residue.

They introduced phosphate.

They extended selected simulations.

And each stage generated another question for the next computational stage.

That is an HPC-driven scientific workflow.

The supercomputer effectively becomes a laboratory in which researchers can repeatedly perturb a molecular system, observe its response, and identify mechanisms that can subsequently be tested experimentally.

The Next Generation of the Calculation

The authors themselves identify where the computational campaign should go next.

Classical MD and endpoint MM-PB(GB)SA can capture conformational sampling and bulk electrostatic effects, but they do not resolve charge transfer, electronic polarization, or chemical bond rearrangement.

The experimentally relevant ribose-donor selectivity involves the Michaelis complex and transition state with phosphate present.

That pushes the problem toward QM/MM or QM-cluster calculations.

At the same time, understanding the kinetics of the conformational transitions will require enhanced-sampling approaches such as:

  • metadynamics;
  • accelerated molecular dynamics;
  • transition-path sampling; and potentially
  • other rare-event sampling techniques.

These approaches are often significantly more computationally intensive than standard trajectory generation, which highlights the critical role of HPC. As scientific inquiries shift from identifying visited enzyme configurations to mapping the transition pathways between them and detailing the chemical reaction mechanisms, the computational demands grow increasingly sophisticated.

The Supercomputing Takeaway

This study does not posit that high-performance computing has fully elucidated the complete mechanism of pyrimidine-nucleoside phosphorylase. Instead, it demonstrates a more significant advancement for computational science: the capacity for HPC to transform static molecular structures into statistically rigorous, sampled computational experiments.

By executing a 156-trajectory campaign, the researchers examined 13 molecular probes across four conformational states with triple replication, yielding 15.6 microseconds of molecular-dynamics data. From this ensemble, the team derived a candidate molecular lid, identified residue-level interaction networks, and generated experimentally testable hypotheses.

Concurrently, the research delineates the inherent boundaries of classical simulation: dynamic retention is not equivalent to equilibrium affinity; computational mutations do not replace empirical laboratory validation; and force-field trajectories do not capture quantum-mechanical reaction pathways. Furthermore, a statistically significant signal does not entirely negate the limitations imposed by the initial structural data.

These distinctions are not deficiencies of HPC; rather, they characterize the iterative nature of modern scientific workflows. In this model, the computer provides molecular evidence, which is then challenged by the scientist and refined through subsequent calculations or alternative computational methods. The broader potential of supercomputing in molecular science lies not merely in increased speed, but in the ability to investigate the complex mobility of molecular machines, a task that remains impossible through the analysis of static structures alone. This study, originating from four protein configurations, concludes with a compelling hypothesis regarding a single tyrosine residue and provides a clear roadmap for future, more sophisticated computational inquiry.

AI agents search 1.9 billion protein clusters, discover a new biological system
Featured

AI agents search 1.9 billion protein clusters, discover a new biological system

Tyler O'Neal, Staff Editor September 25, 2026, 1:00 pm

Autonomous scientific computing turns genome mining into an adaptive HPC workload, using 949 agent sessions, 60 CPU cores, GPU-accelerated structure prediction, and more than 215 million tokens to uncover a previously unknown family of reverse transcriptases.

For decades, a central challenge in computational biology has been deceptively simple to articulate: the volume of biological data far exceeds the capacity for human analysis. While modern metagenomic databases contain billions of uncharacterized protein sequences, conventional computational pipelines remain limited by their reliance on predefined search criteria, effectively restricting discovery to what researchers already know how to describe.

Recent research from Anthropic proposes an alternative paradigm. By moving beyond fixed analytical pipelines, the researchers developed an autonomous system wherein AI agents can search, analyze, self-critique, and initiate iterative computational tasks to investigate unexpected observations. In this study, the system surveyed approximately 1.9 billion protein clusters, ultimately identifying a novel family of reverse transcriptases associated with tandem DNA repeat arrays, which the team has termed array-associated reverse transcriptases (ART).

While this discovery is biologically significant, the underlying computing architecture offers a compelling advancement for high-performance computing (HPC). Rather than a singular AI model addressing a solitary query, the system functioned as an orchestrated scientific workload. Over 21.5 hours of wall-clock time, the campaign executed 119 research tasks and 949 agent sessions, totaling 215.6 million tokens, all without human intervention. The infrastructure utilized 58 concurrent sessions within a sandbox environment featuring 60 CPU cores and 192 GiB of memory, while offloading specialized structure prediction tasks to NVIDIA A100 and L4 GPUs. This experiment illustrates a shift in HPC workloads, moving from conventional batch processing toward dynamic, exploratory scientific computing.

From fixed pipelines to computational exploration

Traditional genome mining is extraordinarily powerful.

Researchers can construct profile hidden Markov models, search enormous sequence databases, cluster homologous proteins, build phylogenetic trees, and examine genomic neighborhoods. These operations are highly amenable to parallel computing.

But there is a fundamental limitation.

A conventional pipeline has to be told what constitutes an interesting result.

If the software is searching for a particular protein family, genomic architecture or sequence motif, the pipeline is optimized around those expectations. Anything that falls outside the predefined feature set may simply be classified as noise.

The researchers behind the study describe this as a novelty problem.

A human scientist looking at a sequence can notice something that was not part of the original search specification: an unusual repeat, a strange genomic neighborhood or an unexpected combination of molecular components.

The question was whether an AI-driven computational system could perform some of that exploratory work at database scale.

The answer, in this experiment, was yes, but with important qualifications.

The system did not simply unleash a language model on 1.9 billion sequences.

It constructed a hierarchy of computational agents.

A launch agent converted the research brief into stages. Worker agents performed individual analyses. Supervisor agents reviewed their plans and results. Curator agents placed findings into a shared knowledge base. Editor agents reviewed reports.

Most importantly, observations could generate new work.

Of the 119 tasks in the campaign, agents proposed 98 follow-up tasks. Those tasks entered a triage queue, where the research harness could release or reject them.

That creates a very different computational model from a conventional workflow.

Instead of: input → fixed pipeline → output the architecture becomes: input → analysis → observation → new task → analysis → new observation → new task

The compute graph can therefore change as the science develops.

For HPC architects, that distinction is crucial.

1.94 billion protein clusters become a compute problem

The initial search was enormous.

The agents assembled reverse-transcriptase profile HMMs and searched approximately 1.94 billion protein clusters.

That produced approximately 198,290 RT clusters after filtering, which were classified into nine RT classes.

The system then examined approximately 10,983 RT loci and evaluated 3,564 recurring protein families in their genomic neighborhoods as potential partner genes.

Sixteen candidate families passed the initial criteria and were assigned dedicated investigations.

An additional candidate emerged from follow-up work.

The campaign ultimately produced 19 reports, including reports on candidate partner families and three newly identified RT lineages.

Only three of the 17 candidate partner families survived as previously unreported RT associations.

That rejection rate is important.

The system was not simply programmed to turn every unusual observation into a discovery. It had to eliminate annotation artifacts, previously characterized systems and proteins that merely happened to occur nearby.

That is where the computational workflow begins to look increasingly like an HPC-enabled scientific laboratory.

The infrastructure behind the agents

The study’s autonomous research harness was built around Claude Code instances configured with Claude Mythos 5.

The computational environment permitted up to 58 concurrent sessions.

The sandbox itself contained:

  • 60 CPU cores
  • 192 GiB of memory
  • No GPU

The absence of GPUs in the main sandbox is itself revealing.

The dominant workload was not neural-network training or large-scale inference performed locally on an accelerator. Much of the work consisted of conventional scientific computing: sequence searches, data manipulation, clustering, alignment, phylogenetic analysis, scripting, file processing and database queries.

The agents could execute software including HMMER, MMseqs2, BLAST+, MAFFT, FastTree, SeqKit, skani, geNomad, Infernal, ViennaRNA and other computational biology tools.

For structural analysis, the workflow could dispatch jobs to external GPU resources.

The researchers report 19 GPU jobs on NVIDIA A100 and L4 GPUs, using ESMFold or ColabFold with AlphaFold2-based models.

That is a classic heterogeneous HPC pattern.

CPU resources handled broad exploration and data analysis.

GPU resources were brought into the workflow when the problem demanded computationally expensive protein-structure prediction.

The AI agents effectively became a workload-management layer sitting above a collection of scientific-computing tools.

The numbers tell the HPC story

The campaign generated:

119 research tasks

949 agent sessions

77 agent-hours

215.6 million tokens

21.5 hours of wall-clock time

7,578 shell commands

696 database queries

131 literature searches

61 web requests

That workload is fundamentally different from a traditional supercomputing simulation.

There is no single enormous MPI job running for several hours across thousands of nodes.

Instead, the workload consists of many relatively independent, heterogeneous and dynamically generated tasks.

Some are computational.

Some are database operations.

Some involve text and literature.

Some involve sequence analysis.

Some invoke GPUs.

Some produce additional work.

And some terminate because the hypothesis is rejected.

This looks less like a conventional batch queue and more like a scientific task graph whose topology is discovered during execution.

That could become an important class of HPC workload.

Then the AI noticed something it wasn’t specifically looking for

The most consequential observation emerged from an investigation that was not originally designed to find the ART system.

The research campaign was primarily looking for previously unknown associations between reverse transcriptases and partner protein-coding genes.

One RT lineage, however, contained an unusual feature in the noncoding DNA upstream of the RT.

An agent retrieved the actual DNA sequence.

And instead of merely processing a predefined annotation, the agent examined the sequence directly.

It noticed a repeating pattern.

The worker identified tandem repeats approximately 16–17 nucleotides long, separated by spacers roughly 100–200 nucleotides long.

One locus contained 14 copies of a 16-nucleotide repeat.

The pattern looked sufficiently unusual that the agent began comparing it with known systems, including CRISPR-like arrays, retron-related architectures and other repeat-associated mechanisms.

It then performed a novelty investigation.

The repeat architecture did not match previously reported features.

That observation became the starting point for the ART discovery.

This is perhaps the most interesting computational moment in the entire study.

The system found something that the original research specification had not explicitly asked it to find.

The discovery emerged from looking at the data, rather than merely matching the data against a predefined list of expected features.

Why context mattered

The researchers subsequently tested whether the AI models could recognize the unusual repeat arrays when given different amounts of information and different tools.

The results expose an important limitation, and an important opportunity.

The strongest models could recognize the ART array when the DNA sequence itself was placed directly into their context.

But giving the model more tools did not automatically make it better at recognizing the repeat architecture.

In some benchmark conditions, performance actually declined.

The researchers found that a major factor was whether the model actually read enough DNA sequence.

When at least 200 nucleotides of contiguous DNA were read into context, repeat recognition increased substantially.

For the pooled group of four most capable models, recognition increased as progressively larger amounts of DNA were brought into context, reaching as high as 76% of attempts in the reported bins and as high as 96% for Mythos 5 in the largest-context condition.

The implication is striking for scientific AI infrastructure.

Giving an agent access to a tool is not the same thing as giving it the information contained in the tool’s output.

A filesystem can contain thousands of nucleotides, millions of rows, or gigabytes of scientific data. The agent still has to decide what to inspect.

That makes data movement, context construction, and intelligent I/O part of the computational problem.

For future scientific AI systems, the bottleneck may not always be FLOPS.

It may be what information the agent chooses to bring into its working context.

From sequence anomaly to biological system

Once the unusual architecture was identified, the researchers expanded the computational investigation.

The ART systems turned out to occur primarily in jumbo phages and to contain three recurring components: an RT protein, a tandem repeat array, and a partner gene.

The researchers examined 95 ART members and additional phage loci.

Computational analysis showed that the repeat arrays were not simply random sequence repetitions. Their spacing, conservation and organization distinguished them from shuffled controls.

Phylogenetic analyses placed the ART proteins in a distinct family.

The researchers also identified three major partner-protein types.

Type I systems were associated with a protein of roughly 600 amino acids containing two tandem GNAT-like folds.

Type II systems occurred in a Staphylococcus phage lineage and encoded an approximately 270-residue all-helical protein.

Type III systems encoded a smaller, approximately 170-residue helical protein in one environmental lineage.

The three partner families showed no obvious sequence or predicted-structure similarity to one another.

That suggests the RT family may have been paired with unrelated partner proteins multiple times during evolution.

Again, much of this characterization depended on computational infrastructure.

Sequence clustering and homology searches narrowed the candidate universe.

Multiple sequence alignment and phylogenetic tools established evolutionary relationships.

Structure-prediction systems provided additional evidence about protein architecture and possible RT-partner interfaces.

The CPU and GPU workloads were therefore not competing alternatives.

They were complementary stages of the same scientific pipeline.

The GPU was not the discovery engine, it was part of the investigation

This distinction is worth emphasizing.

It would be easy to describe the work as an AI supercomputer discovering a new molecular system.

That would obscure how the computation actually worked.

The initial genome-scale search relied heavily on sequence analysis tools and CPU resources.

The GPU resources were used for specialized structural prediction.

That division is representative of the emerging architecture of scientific AI.

The future supercomputing system may not be dominated by a single accelerator type.

Instead, a scientific workload could move repeatedly between: CPU search → database → agent reasoning → CPU analysis → GPU prediction → CPU comparison → agent interpretation → new task

The scheduler becomes more important because the application itself determines what it needs next.

The workload is no longer completely known before execution begins.

The shared knowledge base becomes a new kind of scientific memory

Another important architectural feature was the shared knowledge base.

After each task, a curator agent reviewed the work and entered findings into a common record.

Later agents received relevant entries in their prompts.

Plans, results, reviews, and scripts were also maintained in a version-controlled record accessible to the agents.

This effectively created a persistent computational memory for the research campaign.

In HPC terms, it resembles a combination of workflow state, provenance database, experiment log and shared scientific scratch space, but with the contents actively influencing future computation.

That is potentially a major architectural direction for agentic science.

A traditional HPC workflow generally has explicit inputs and outputs.

An agentic workflow needs something more.

It needs to remember:

  • what has already been tested;
  • which hypotheses failed;
  • which datasets produced useful evidence;
  • which computational tools were successful; 
  • which observations deserve follow-up;
  • what another agent has already learned.

Without that shared state, dozens or hundreds of agents would simply duplicate each other’s work.

The database therefore becomes part of the intelligence of the system.

A supercomputer that changes the question

Perhaps the most inspiring aspect of the study is not that AI found another protein family.

It is that the computational system changed the question being asked.

The original mission focused on RT partner genes.

The important discovery emerged from noncoding DNA.

An agent saw a pattern that existed outside the original feature specification and created a new investigative path.

That is a subtle but potentially profound change in scientific computing.

For generations, supercomputers have excelled at executing questions defined by humans.

Scientists formulate a model.

They construct the equations.

They select the parameters.

They define the search space.

The machine explores that space at extraordinary speed.

Agentic scientific computing suggests another model: the machine can help explore what the search space should have been.

That does not mean the machine becomes the scientist.

It means the computational system can participate in identifying anomalies that deserve human attention.

And that distinction matters.

The discovery still requires science beyond the computer

The study should not be interpreted as proof that an AI independently solved the biological function of ART.

The computational evidence is substantial, but the biological mechanism remains incompletely understood.

The researchers observed that ART arrays produce discrete RNAs and that these RNAs can be highly expressed during phage infection. They also observed corresponding RNA production when ART systems were expressed in E. coli.

Those observations support the hypothesis that the arrays generate a repertoire of RNA molecules.

But the exact biological function of the ART system remains an open question.

The researchers have not established the complete biochemical mechanism by computational analysis alone.

That is where laboratory experimentation remains essential.

This is an important boundary for autonomous scientific computing.

AI can search.

AI can classify.

AI can notice anomalies.

AI can generate hypotheses.

AI can design follow-up analyses.

But the distinction between a compelling computational hypothesis and an experimentally established biological mechanism remains fundamental.

The next HPC workload may be adaptive

For supercomputing centers, the ART study points toward a workload category that could become increasingly common.

Scientific computing has traditionally been organized around relatively predictable workloads.

A researcher submits a simulation.

A scheduler allocates resources.

The computation runs.

Results are returned.

Agentic science introduces a feedback loop.

A computation produces an observation.

The observation changes the next computation.

The next computation may require a different resource.

A CPU-intensive search might trigger a GPU structure prediction.

The structure prediction might trigger another sequence search.

That result might launch a literature search.

The literature search might trigger a new biological hypothesis.

The hypothesis might generate dozens of additional jobs.

The workload becomes adaptive rather than predetermined.

That creates difficult problems for HPC infrastructure.

Schedulers will need to deal with bursts of short-lived tasks alongside traditional large jobs.

Resource managers may need to coordinate CPU, GPU, memory, and storage allocations dynamically.

Workflow systems will need robust checkpointing and provenance.

Data-management systems will need to move information rapidly between persistent databases, compute nodes and AI context windows.

And scientific users will need ways to reproduce an agent’s decisions, not simply reproduce the final executable.

From FLOPS to scientific decisions

For years, supercomputing performance has been measured in familiar units: FLOPS, bandwidth, latency, memory capacity and energy efficiency.

Those metrics remain essential.

But autonomous scientific computing introduces another dimension.

How efficiently can a machine decide what computation should happen next?

That is a very different performance question.

The ART campaign processed nearly two billion protein clusters, but the important computational achievement was not brute-force enumeration alone.

It was the ability to progressively reduce that enormous search space while retaining the possibility of following an unexpected clue.

The system moved from approximately 1.94 billion clusters to roughly 198,000 RT clusters, then to approximately 11,000 loci, thousands of candidate partner families, and ultimately a much smaller set of biological systems worthy of deep investigation.

The computational hierarchy became a scientific funnel.

At every stage, compute reduced uncertainty.

And occasionally, an unexpected observation widened the funnel again.

That is precisely what makes the workload interesting for HPC.

The beginning of autonomous discovery infrastructure

The study regarding array-associated reverse transcriptases (ART) does not signify the obsolescence of conventional supercomputing; rather, it represents a pivotal transition toward a new paradigm of scientific machinery. The supercomputer of the future may evolve beyond merely accelerating simulations to orchestrating thousands of heterogeneous computational operations. Such a system would maintain a collective memory of scientific evidence, determine the necessity of further computation, and dynamically route tasks to the optimal hardware.

In this model, CPU cores would manage genome searches, GPUs would facilitate molecular structure prediction, and specialized storage and networking would handle vast sequence repositories and intermediate datasets. AI agents would serve as the decision-making layer, identifying which investigations warrant further resources, while human scientists remain at the forefront to validate emerging discoveries. The ART discovery offers a preliminary look at this architecture, demonstrating how a vast database can be transformed into an active search space and how a suite of diverse tools can function as a unified scientific instrument. Ultimately, the next generation of supercomputing will likely transcend simple calculation, increasingly assisting researchers in discerning which scientific questions are truly worth pursuing.

AI’s trillion dollar compute race hits a hard limit: There isn’t enough power
Featured

AI’s trillion dollar compute race hits a hard limit: There isn’t enough power

Tyler O'Neal, Staff Editor September 24, 2026, 10:00 am

For years, the artificial intelligence industry has focused on a deceptively simple question: how many GPUs can be integrated into a single data center? However, an increasingly critical challenge has emerged that can no longer be ignored: securing the electricity required to power these systems. This question is now fundamentally shaping the future of supercomputing.

The urgency of this issue is highlighted by Oracle’s Project Jupiter in New Mexico, a cornerstone of the collaborative Stargate infrastructure initiative involving Oracle, OpenAI, and SoftBank. Oracle has issued a force-majeure notice to the project’s developer, a unit of Blue Owl Capital, citing potential delays in securing sufficient power. While this notice provides contractual flexibility should the 2028 operational target be missed, Oracle maintains that the project remains on schedule.

The implications for the broader supercomputing industry extend far beyond a single facility or financing arrangement. This situation reveals a fundamental systemic risk within the current AI infrastructure boom: it is becoming significantly easier to acquire computational capacity than to secure the physical infrastructure necessary to power it.

The supercomputer is no longer just a computer

Traditional supercomputing discussions tend to revolve around familiar metrics: FLOPS, accelerator count, memory bandwidth, network bandwidth, storage throughput, and application performance.

Those metrics remain critical.

But an AI supercomputer also has another specification that is becoming just as important: Megawatts.

Modern AI clusters are effectively enormous distributed computing systems. Thousands of accelerators must operate simultaneously, connected by extremely high-bandwidth networks and supported by storage, cooling, and power-conversion infrastructure.

The result is a system whose computational performance is inseparable from its physical infrastructure.

A facility may have the latest accelerators available.

It may have the network fabric.

It may have the cooling system.

It may even have customers waiting for compute capacity.

But if the electrical infrastructure is not ready, the supercomputer does not exist in any meaningful operational sense.

It is simply an expensive collection of hardware waiting for electrons.

Project Jupiter Makes the Problem Concrete

Project Jupiter illustrates the scale of the challenge.

The New Mexico campus is designed as a massive AI computing facility. Recent reporting puts its planned power requirement at roughly 2.45 gigawatts, with the current design centered on Bloom Energy fuel cells operating as an onsite microgrid. 

That is not a conventional data-center power requirement.

It is an industrial-scale energy system attached to a computing system.

And the power infrastructure itself has become a critical-path component.

A natural-gas pipeline intended to supply the facility has faced regulatory setbacks and a delay. TechCrunch reports that the pipeline schedule has moved to February 2027, while a separate air-quality permit for the fuel-cell system remains pending. 

Oracle’s own June description of the revised design says the company moved away from the previously planned gas-turbine and diesel-generator configuration toward Bloom Energy fuel-cell technology. Oracle says the revised system is intended to reduce water consumption and nitrogen-oxide emissions while providing reliable onsite power. 

That engineering evolution is important.

It also demonstrates the uncomfortable reality of AI infrastructure: The power system can become as complicated as the computer system.

When megawatts become a computing specification

Consider what happens inside a large AI cluster.

An accelerator performing a computation consumes electrical power.

Thousands of accelerators multiply that requirement.

Then add CPUs, memory systems, high-speed networking, storage, power-conversion losses, cooling equipment and facility overhead.

The electricity requirement becomes enormous.

And unlike purchasing additional GPUs, increasing electrical capacity is not simply a matter of placing another order.

Power infrastructure requires physical construction.

Transmission capacity may have to be expanded. Substations must be built. Generation resources have to be secured. Fuel infrastructure may be required. Permits have to be obtained. Cooling systems must be engineered. Communities and regulators may have to approve the development.

Those processes operate on very different timescales from the semiconductor industry.

A new accelerator generation can arrive in months.

A major power project can take years.

That mismatch is becoming one of the central infrastructure problems of the AI era.

The GPU supply chain may not be the only bottleneck

The technology industry has spent enormous resources expanding accelerator production.

That effort has created another race: the race to build facilities capable of deploying those accelerators at scale.

This changes the economics of supercomputing.

If a company can acquire 100,000 accelerators but cannot energize the corresponding computing facility, those accelerators do not produce useful AI capacity.

The limiting resource has shifted from silicon alone to the entire infrastructure stack.

Compute availability = accelerators + memory + networking + storage + cooling + power + facility.

Remove any one of those components and the system’s theoretical performance becomes irrelevant.

For AI infrastructure developers, this creates a dangerous possibility: billions of dollars can be committed to computational capacity before the physical infrastructure necessary to operate that capacity is fully secured.

Project Jupiter demonstrates precisely why that matters.

The financing problem follows the power problem

There is another layer to this story.

The AI infrastructure boom is being financed at a scale rarely seen in computing.

Project Jupiter reportedly has approximately $18 billion in loans tied to its development, while Blue Owl has committed roughly $3 billion in equity to the New Mexico project, according to reporting from The Information. 

Reuters reported last week that the $18 billion in loans had come under pressure, with portions quoted around 89 to 91 cents on the dollar amid concerns about the project’s regulatory and infrastructure challenges. 

That does not mean the project has failed.

It does, however, demonstrate how the physical risks of AI infrastructure can become financial risks.

If a supercomputer takes longer than expected to come online, capital remains tied up.

If power infrastructure is delayed, the facility cannot generate the expected computing capacity.

If construction costs rise, financing requirements increase.

If customer commitments depend upon a particular operational date, delays can ripple through the entire AI infrastructure ecosystem.

The computer may be digital.

The risk is not.

The AI factory has become an energy factory

There is a conceptual shift taking place in the industry.

The next generation of AI facilities should perhaps no longer be thought of simply as data centers.

They are AI factories.

They convert electricity into computation.

Electricity enters the facility.

Accelerators transform that energy into mathematical operations.

Networks move data between processors.

Memory systems feed the calculations.

Storage provides the datasets.

Cooling removes the resulting heat.

The output is computational capacity.

From that perspective, electricity is not merely an operating expense.

It is one of the fundamental raw materials of AI.

That makes the availability of electricity a direct determinant of how much AI computation a company can actually deliver.

Efficiency suddenly matters more

This also changes the meaning of performance optimization.

Historically, HPC engineers have pursued better performance for familiar reasons: finish the simulation sooner, increase throughput, reduce queue times or solve larger problems.

AI adds another dimension: How much computation can be produced per megawatt?

That question could increasingly influence processor architecture, cooling technology, interconnect design, scheduling software, and even algorithms.

A cluster that delivers more useful work per watt can effectively provide more computational capacity without requiring proportional increases in generation and transmission infrastructure.

This is where traditional HPC engineering becomes particularly relevant.

Techniques developed to maximize utilization of supercomputers, workload scheduling, accelerator efficiency, communication optimization, memory locality, precision reduction, and application-specific optimization, suddenly have an infrastructure-level economic consequence.

Every percentage point of efficiency can represent substantial avoided power consumption when multiplied across hundreds of megawatts.

The hidden supercomputer bottleneck

The industry has become accustomed to thinking about AI bottlenecks in terms of GPUs.

Then came high-bandwidth memory.

Then networking.

Then advanced packaging.

Now another bottleneck is becoming increasingly visible: The grid.

Project Jupiter is not proof that the AI industry has run out of electricity.

It is evidence that obtaining enough reliable power, in the right location and on the required schedule, is becoming a major engineering and infrastructure challenge for hyperscale AI.

That distinction matters.

Oracle maintains that Project Jupiter remains on schedule, and the company has invested heavily in a revised onsite power strategy. Oracle also says it will fund the project’s energy infrastructure and electricity costs rather than shifting those costs to local residents. 

But the fact that power availability has become important enough to appear in a force majeure notice should get the attention of anyone planning the next generation of AI supercomputers.

The Clock Is Running

There is an uncomfortable mismatch at the heart of the AI boom.

The semiconductor industry is accelerating.

AI models are growing.

Demand for inference is expanding.

Training clusters are becoming larger.

Hyperscalers are announcing increasingly ambitious AI infrastructure programs.

But electrical infrastructure cannot necessarily move at the same speed.

The industry can announce a gigawatt-scale AI campus today.

That does not mean the electrons will be available tomorrow.

And without those electrons, the promised FLOPS remain theoretical.

This is why Project Jupiter deserves attention from the supercomputing community.

The story is not fundamentally about Oracle’s stock price, Blue Owl’s investment or one delayed pipeline.

It is about whether the physical infrastructure of the world’s computing systems can keep pace with the computational ambitions of the AI industry.

The next supercomputing race may be measured in megawatts

For decades, progress in supercomputing was primarily measured in FLOPS. Over time, the industry’s focus expanded to include memory bandwidth, interconnect performance, storage throughput, and energy efficiency. Currently, however, a critical new metric has emerged: available power. The next generation of supercomputing facilities may be constrained not by the density of processors, but by the volume of megawatts that can be reliably delivered to the site. This introduces a significant uncertainty into the trillion-dollar AI infrastructure race. 

While the industry may possess sufficient chips, capital, customers, and data, these assets remain dormant without the necessary electricity to power them. Ultimately, the future of artificial intelligence may depend on a fundamental infrastructure challenge: whether we can scale power generation in alignment with our computational ambitions.

  • Supercomputing turns dark-matter waves into a testable prediction
  • 1
  • 2
Page 1 of 2
POPULAR RIGHT NOW
  • Supercomputers reveal four regimes of radiation damage in tungsten
    Supercomputers reveal four regimes of radiation damage in tungsten
  • Supercomputing rewrites the timeline of planet formation at cosmic dawn
    Supercomputing rewrites the timeline of planet formation at cosmic dawn
  • The next supercomputing breakthrough may come from memory, not compute
    The next supercomputing breakthrough may come from memory, not compute
  • 10,000 AI agents, 130 billion tokens and 88 hours: How OpenAI turned Navier–Stokes into a supercomputing workload
    10,000 AI agents, 130 billion tokens and 88 hours: How OpenAI turned Navier–Stokes into a supercomputing workload
  • Supercomputing reveals why some black hole flares fade away
    Supercomputing reveals why some black hole flares fade away
  • NVIDIA's $96.2 billion quarter redefines the supercomputing economy
    NVIDIA's $96.2 billion quarter redefines the supercomputing economy
  • When physics computes: Simulations turn random skyrmion motion into directional information
    When physics computes: Simulations turn random skyrmion motion into directional information
  • Millions of CPU cores meet 69 billion molecules: AI rewrites the rules of computational drug discovery
    Millions of CPU cores meet 69 billion molecules: AI rewrites the rules of computational drug discovery
  • Jensen Huang to G20: Build the AI infrastructure, or risk being left behind
    Jensen Huang to G20: Build the AI infrastructure, or risk being left behind
  • Milky Way’s own gravity can mimic dark matter clues, supercomputer simulations suggest
    Milky Way’s own gravity can mimic dark matter clues, supercomputer simulations suggest
THIS YEAR'S MOST READ
  • Wall Street wants to trade supercomputing power like oil
    Wall Street wants to trade supercomputing power like oil
  • Beamforming the future: BeammWave's 6G push signals the rise of orbital-terrestrial wireless networks
    Joakim Axmon
    Joakim Axmon
  • AI breaks conservation barriers: Australia’s Wildlife Observatory leverages supercomputing to protect biodiversity
    AI breaks conservation barriers: Australia’s Wildlife Observatory leverages supercomputing to protect biodiversity
  • Silicon spintronics brings the P-computer closer to reality
    Microscope image of a semiconductor-integrated spintronic test chip developed by researchers at Tohoku University and NIST. The device demonstrates the first silicon-integrated probabilistic bit (p-bit), a key building block for future large-scale probabilistic computers designed for AI and optimization workloads.
    Microscope image of a semiconductor-integrated spintronic test chip developed by researchers at Tohoku University and NIST. The device demonstrates the first silicon-integrated probabilistic bit (p-bit), a key building block for future large-scale probabilistic computers designed for AI and optimization workloads.
  • Physics-trained ‘Digital Super Brain’ learns from supercomputers to accelerate discovery
    Physics-trained ‘Digital Super Brain’ learns from supercomputers to accelerate discovery
  • Hidden order, revealed at scale: Supercomputing, electron ptychography uncover the inner workings of relaxor ferroelectrics
    Hidden order, revealed at scale: Supercomputing, electron ptychography uncover the inner workings of relaxor ferroelectrics
  • Cosmic ambition at scale: UK’s supercomputer unlocks a 2.5 petabytes universe
    Cosmic ambition at scale: UK’s supercomputer unlocks a 2.5 petabytes universe
  • Memory has become the new compute: Why Micron, SK Hynix crossing $1 trillion matters to supercomputing
    Memory has become the new compute: Why Micron, SK Hynix crossing $1 trillion matters to supercomputing
  • Huawei’s Tau Scaling ambition tests the limits of post-Moore semiconductor reality
    He Tingbo from HUAWEI delivered a keynote speech titled "New Semiconductor Path in Practice"
    He Tingbo from HUAWEI delivered a keynote speech titled "New Semiconductor Path in Practice"
  • Intel's Q1 results signal supercomputing surge driving Xeon momentum
    Intel's Q1 results signal supercomputing surge driving Xeon momentum
MOST READ OF ALL-TIME
  • Largest Computational Biology Simulation Mimics The Ribosome
    Details
    112550
    The amino acid (green) slithers into the chemical reaction center, moving through an evolutionarily ancient corridor of the ribosome (purple). The amino acid is delivered to the reaction core by the transfer RNA molecule (yellow).
    The amino acid (green) slithers into the chemical reaction center, moving through an evolutionarily ancient corridor of the ribosome (purple). The amino acid is delivered to the reaction core by the transfer RNA molecule (yellow).
  • Silicon 'neurons' may add a new dimension to chips
    Details
    81852
    Silicon 'neurons' may add a new dimension to chips
  • Linux Networx Accelerators Expected to Drive up to 4x Price/Performance
    Details
    76035
  • Complex Concepts That Really Add Up
    Details
    74420
    Complex Concepts That Really Add Up
  • Blue Sky Studios Donates Animation SuperComputer to Wesleyan
    Details
    68599
    Each rack holds 52 Angstrom Microsystem-brand “blades,” with a memory footprint of 12 or 24 gigabytes each. (Photos by Olivia Bartlett Drake)
    Each rack holds 52 Angstrom Microsystem-brand “blades,” with a memory footprint of 12 or 24 gigabytes each. (Photos by Olivia Bartlett Drake)
  • Humanities, HPC connect at NERSC
    Details
    58458
  • TeraGrid ’09 'Call for Participation'
    Details
    55439
  • Turbulence responsible for black holes' balancing act
    Details
    52822
  • Cray Wins $52 Million SuperComputer Contract
    Details
    50594
  • SDSC Researchers Accurately Predict Protein Docking
    Details
    46668
  • FRONTPAGE
  • LATEST
  • POPULAR
  • REGISTER
  • SOCIAL
  • VIDEO
  • SUBSCRIPTION
  • RSS
  • GUIDELINES
  • PRIVACY
  • TOS
  • ABOUT
  • +1 (816) 799-4488
  • editorial@supercomputingonline.com
© 2001 - 2026 SuperComputingOnline.com, LLC. All rights reserved. This material may not be published, broadcast, rewritten or redistributed without permission.
Sign In
  • FRONT PAGE
  • LATEST
    • MEDIA KIT
    • MOST READ
    • RSS FEED
    • ACADEMIA
    • AEROSPACE
    • APPLICATIONS
    • ASTRONOMY
    • AUTOMOTIVE
    • BIG DATA
    • BIOLOGY
    • CHEMISTRY
    • CLIENTS
    • CLOUD
    • DEFENSE
    • DEVELOPER TOOLS
    • EARTH SCIENCES
    • ECONOMICS
    • ENGINEERING
    • ENTERTAINMENT
    • HEALTH
    • INDUSTRY
    • INTERCONNECTS
    • GAMING
    • GOVERNMENT
    • MANUFACTURING
    • MIDDLEWARE
    • MOVIES
    • NETWORKS
    • OIL & GAS
    • PHYSICS
    • PROCESSORS
    • RETAIL
    • SCIENCE
    • STORAGE
    • SYSTEMS
    • VISUALIZATION
  • VIDEOS
    • ADD YOUR VIDEOS
    • MANAGE VIDEOS
  • COMMUNITY
    • TRADE SHOWS
    • SOCIAL NETWORK VIDEOS
    • SURVEYS
    • APPLICATIONS BROWSER
    • CONVERSATION INBOX
    • SOCIAL ADVERTISER
    • GROUPS
    • MARKETPLACE LISTINGS
    • PAGES
    • LEADERBOARD
    • POINTS LISTING
      • BADGES
    • PRIVACY CONFIRM REQUEST
    • PRIVACY CREATE REQUEST

Hey there! We noticed you’re using an ad blocker.