Friday, October 2, 2026

Inside NVIDIA's 113-Degree Cooling Design, and Westfield's Data Center Question

Given the strong debate and varied perspectives surrounding the proposed Westfield (my hometown) data center campus, this post focuses strictly on verifiable technical specifications and empirical data as I understand it. The goal is to provide an objective, fact-based overview of the project's engineering parameters, NVIDIA's recent high-temperature liquid cooling design, local climate limits, and broader infrastructure trade-offs. Read my earlier post on the topic here.

I grew up at 342 Holyoke Road in Westfield, Massachusetts. In 2021, the City Council approved a $4 billion data center campus at 199 Servistar Industrial Way, 1.8 miles from that house. The vote passed 9 to 3. The site sits partly on wetlands and directly over the Barnes Aquifer, the region's municipal drinking water source. In July 2026, the council voted unanimously for a one year moratorium on new data centers, and the campus would draw 274 megawatts. Servistar says it plans a closed-loop cooling system. The reporting I reviewed does not name the cooling vendor or design.

In June 2026, NVIDIA published a cooling design that addresses the same question. It appears in the DSX reference design, a guide for building the full AI factory infrastructure stack. NVIDIA's Ali Heydari says the design has zero water consumption and has eliminated most power use for cooling. NVIDIA's own description of the Vera Rubin NVL72 system calls it single-phase direct liquid cooling with a 113°F supply temperature.

Earlier liquid-cooled deployments supplied water in the 80°F to 90°F range. Those systems cooled the CPUs and GPUs with cold plates and left other components to air. Rubin cools every chip and networking component by liquid, eliminating fans inside the server chassis and compute racks. The coolant is a 75 percent water and 25 percent propylene glycol mix that enters the chip at 113°F and leaves near 131°F. Operators have traditionally recommended an ambient temperature of 64°F to 81°F. NVIDIA's summary also lists a 6U system now fitting in 2U and no hot or cold aisle management.

Inside the building, a Cooling Distribution Unit (CDU) uses a liquid-to-liquid heat exchanger to isolate the secondary coolant loop in the server racks from the primary facility loop. The primary loop then carries heat outside to dry coolers. NVIDIA says they reject heat efficiently for much of the year and that the loop avoids evaporative cooling 99 percent of the time. In favorable climates, NVIDIA says this cuts water use from roughly 2.6 million gallons per megawatt per year to near zero. Hotter climates such as Phoenix may still need chillers on peak summer days. The captured heat can also be reused to warm nearby buildings.

Heat path from chip to outdoor air. Coolant temperatures are NVIDIA figures.

Westfield's climate aligns with the thermal operating requirements for those dry coolers. The average July high is 83°F, and the average January high is 34°F. Barnes Municipal Airport recorded a high of 98.1°F in July 2026. Because fluid leaves the chips near 131°F and typical heat exchangers require a 5°F to 9°F approach margin, outdoor air only needs to remain below about 104°F to supply 113°F coolant back to the racks. As a result, dry coolers in Westfield can maintain a 113°F supply temperature year-round without requiring supplemental mechanical chillers. Nearby Springfield has gained 11 more above-average summer days since 1970, but local peaks remain within the system's thermal operating limit. A campus drawing 274 megawatts releases close to that much heat, and dry coolers send it to outdoor air.

Westfield outdoor temperatures compared with coolant temperatures. The 104°F limit is an estimate.

NVIDIA estimates that a 50MW facility saves over $4 million a year in cooling energy and water. Cooling has accounted for up to 40 percent of a data center's electricity. One reviewer notes the capital cost premium over air cooling remains unknown.

Environmental and energy analyses point out that on-site claims exclude water consumed by off-site power plants supplying the grid. Lawrence Berkeley National Laboratory found that 92.5 percent of a data center's water footprint comes from generating its electricity. Dry coolers can also demand 10 to 35 percent more electricity than evaporative towers. Servistar's plans include natural gas generators, and Westfield Gas & Electric estimates a gas pipeline at about $20 million.

State certification standards will require the developer to disclose cooling method, power source, and noise mitigation before construction proceeds. 

Every source is linked above for you to read and weigh.

Wednesday, September 30, 2026

IonQ Runs Real-Time Error Decoding on a Single CPU

On September 22, IonQ reported a real-time error decoder that runs on one standard off-the-shelf CPU. Decoding is the classical half of quantum error correction, and it sets a speed limit on the quantum half. This post covers what the decoder does, the architecture it serves, and what the test did and did not show.

What a Decoder Does

An error-correcting code measures parity checks on the qubits every cycle. The results, called syndromes, do not reveal the stored data. They show where errors likely occurred. A classical decoder reads the syndrome stream, infers the most probable errors, and hands corrections back to the machine. Some logical operations wait on that answer before the program can continue, so a slow decoder stalls the quantum computer. IonQ describes this as the failure of conventional approaches, where the classical side gets overwhelmed and the quantum system has to pause. Streaming decoders address it by processing syndromes faster than they accumulate.

The Architecture It Serves

IonQ published its Walking Cat architecture in April as a full blueprint for a trapped-ion fault-tolerant computer, covering the compiler, error-correction protocols, micro-architecture, and decoder. It builds entirely on low-density parity-check (LDPC) codes. In the notation [[n, k, d]], n is physical qubits, k is logical qubits, and d is the code distance. The paper introduces a [[70, 6, 9]] code for fast logical gates and a [[102, 22, 9]] code that packs 22 logical qubits into each memory block.

The hardware is a quantum charge-coupled device (QCCD) chip. Electric fields shuttle ions between storage and interaction zones, which lets the chip implement the non-local connections LDPC codes need. A cat factory produces cat states that travel through the machine and get consumed by logical operations. Reservoirs of fresh ions replace qubits lost during operation. The paper's dense design reaches 110 logical qubits and about one million T gates per day with 2,514 physical qubits. Speed is the tradeoff: IonQ estimates a 30-bit Shor factoring run takes about 23 hours.

Trapped-ion cycles run slower than superconducting ones, so the decoder works on a millisecond-scale budget. The April paper described a streaming beam decoder that works on syndrome data in sliding windows to fit that budget.

Testing the Decoder at 408 Logical Qubits

The new paper, Real-time decoder for a MegaQuOp quantum computer using a single CPU, evaluates a dual-decoder architecture on benchmark circuits. The circuits simulate up to 408 logical qubits across 88 memory blocks and magic state factories, and they run more than 31.5 million operations. Under standard operational noise, the decoder added as little as 0.02% stretch time, meaning extra run time spent waiting on decoding. The simulation exceeds the 110 logical qubits in the April dense design.

IonQ says the result shows classical hardware overhead does not have to grow exponentially with logical qubits or circuit depth. That claim matters for scaling, because a decoder that needs more classical hardware with every added qubit would cap the machine size. The company's roadmap runs past 256 physical qubits toward thousands.

Limits of the Result

Every circuit was simulated. The figures come from IonQ's own paper and press release, and they reflect a standard noise model. Hardware adds its own error mix, including the ion loss the April architecture handles with reservoirs. Whether the decoder holds its 0.02% overhead on a running machine is a hardware question, and no device with 408 logical qubits exists to answer it.

For readers of Quantum from the Ground Up: this adds a CPU-based decoder alongside the NVIDIA decoder in Chapter 12, and a classical-side entry to IonQ's coverage in Chapter 6. The two decoders report different metrics, so no ranking follows. The Q-Day range in Chapter 13 stays where it is until decoder results come from hardware. This post will be incorporated into the next edition. The current edition is at gordostuff.com/p/quantum-from-ground-up-hardware.html.

Tuesday, September 29, 2026

950 Claude Agents Found Something in Phage DNA

PCR (Polymerase Chain Reaction) was still something being figured out when I was an undergrad studying microbiology at UMass Amherst. ELISA (Enzyme-Linked Immunosorbent Assay) had been out for a few years but very few labs were running it. . Most of our work still ran on overnight cultures: streak a plate, incubate it overnight, read the colonies in the morning. A literature search meant pulling bound journals off library shelves.

On September 23, Anthropic announced that Claude agents found a previously uncharacterized enzyme system in bacteriophages, the viruses that infect bacteria. Anthropic scientists gave Claude one prompt: search a massive DNA sequence database for interesting reverse transcriptases, enzymes that copy RNA into DNA. About 950 agents ran for 21 hours and used 210 million tokens. The output:

•       200,000+ reverse transcriptases gathered

•       3,500 new candidate systems flagged

•       20 candidates written up as human-readable reports

One agent reading raw DNA beside an unusual enzyme spotted a tandem repeat array by eye. It counted the repeats, measured their spacing, compared the layout with known systems, and checked the literature before filing a report. Anthropic named the system Array-associated Reverse Transcriptases, or ART: the enzyme, a partner gene, and a repeat array. Lab work confirmed the array is transcribed into short RNAs.

Earlier studies had identified the enzyme; Claude was first to connect the array and partner protein. Nobody knows what ART does. The preprint has no peer review yet, and reruns struggled to reproduce the find. CRISPR pioneer Feng Zhang reviewed the preprint and said the result merits further investigation.

Anthropic says this kind of genome mining takes an expert weeks to months. Humans still run every experiment in its lab. The constraint now sits at the bench.

PCR took years to move from a new method into every teaching lab, and it needed Taq polymerase from a Yellowstone hot spring bacterium to get there. ART has 21 hours of agent time and one lab result behind it. Robots can pipette these days but so far anyways, they cannot dream up what to test next…. At least not very well.

950 agents turned a search that would take an expert months into a Tuesday afternoon, and handed a lab twenty leads worth checking.

Thursday, September 24, 2026

DOE Puts $215 Million on 100 Logical Qubits

My quarterly-updated book, Quantum from the Ground Up, tracks the hardware race toward fault-tolerant quantum computers. The Department of Energy just put a date on it. On September 17, DOE launched the Quantum Genesis Q Competition with up to $215 million planned. Teams must deploy at least 100 logical qubits running hundreds of millions of fault-tolerant operations on chemistry, materials, physics, and applied mathematics problems. The evaluation is set for September 2028.

I've written prior - a logical qubit combines many physical qubits under error correction so the result survives a long calculation. Chapter 4 of the book walks through how that works. Some chips already carry more than 1,000 physical qubits, and those qubits still make too many errors to finish useful work. Here's the DOE competition breakdown.

Phase I awards up to $1.5 million per team for early milestones. Phase II holds $100 million for teams that reach 100 logical qubits, plus two $50 million bonus pools at 150 and 200. HPCwire reports that each pool splits equally among the teams that qualify, so a team at 200 logical qubits draws from all three.

       Structure of the DOE Quantum Genesis Q Competition. Source: U.S. DOE, HPCwire

Verification sits with the national labs. DOE plans a $45 million testbed to benchmark competitors' claims independently. Vendors submit results, and the labs check them before any pool pays out.

QuEra and Harvard have shown 96 logical qubits on neutral atoms, the figure cited in Chapter 8, so at least one platform sits near the count. The operation depth and outside verification set the higher bar. Microsoft's Majorana roadmap, covered in Chapter 10, targets 2029, a year after DOE's evaluation.

Two caveats. Only $2.5 million of the $215 million comes from FY2026 funds, and Congress controls the rest. And 100 logical qubits sits far below the roughly 1 million physical qubits Gidney estimated in 2025 for breaking RSA-2048, the baseline in Chapter 13. A 100-logical-qubit machine cannot break RSA-2048, so the schedule for moving to post-quantum encryption stays the same.

Applications close October 19. DOE holds an applicant webinar September 25 at 1:00 p.m. ET. This post goes into the next edition (December 2026) of Quantum from the Ground Up, available here.

Sunday, September 20, 2026

Now That I Know What They Look Like

I spot them now. A small dark box on a pole, angled at the road, above an intersection I have driven through for years. I never noticed it. Once you know the shape, you find them.

They are Flock Safety license plate readers, and a WIRED and 404 Media investigation now documents what one contains. A hacker collective calling itself stegan0gram removed a camera from above a roadway, made a near-complete copy of its storage, and shared the files with both outlets through Distributed Denial of Secrets. As an engineer, I read the technical findings first.

The device runs Android on a processor similar to those in midrange smartphones. About 20 Flock-built apps handle motion detection, image capture, object classification, uploads, and remote updates. When something moves into view, the camera takes a rapid series of photos. A typical vehicle generated about 28 images, and some generated more than 100. The camera varies exposure to capture both the plate and the wider scene, selects and crops useful frames, and sends them over the cellular network. Plate reading and identification of make, model, and color appear to happen on Flock servers.

The recovered logs cover about 21 days across several periods. In that time the camera photographed roughly 50,200 vehicles and generated about 1.6 million images. A typical day logged around 3,300 vehicles, with a high of 4,454. Older logs had been overwritten.

Flock describes the system as protected by on-device encryption. The hackers found two partitions that were unencrypted, named "vendor" and "media." The media partition held an encryption key that unlocked another part of the storage, including videos of thousands of vehicle detections. Much of the most sensitive storage stayed encrypted. In early 2025, researcher Jon Gaines documented root-level flaws in a Flock reader. Flock acknowledged them, said they required physical access, and said an attacker still could not reach footage because images remain on the device only briefly. The hackers had physical access. Flock says it received no report through its vulnerability disclosure process and lacks enough detail to assess their claims.

The software also detects people. It records where each person appears in the image and a confidence score. WIRED extracted the models and ran them against 27,321 stored clips, each one to two seconds long, 1,024 by 768 pixels, with no audio. The models found people in 11 clips, all of them motorcycle riders. The camera points down at traffic, so pedestrians rarely enter the frame. The plate detector also cropped bumper stickers, dealership frames, and an American flag patch on a saddlebag as if each were a plate. Investigators found no evidence of active face recognition, which matches Flock’s statement.

The logs show the hardware under strain: more than 27,000 "no space left on device" errors, plus tens of thousands of related errors, crashes, and reboots. A health check ran about every two minutes and logged "Who’s a good boy?!" more than 12,000 times.

I've taught embedded design courses, and the encryption finding fits a single lecture point. A key stored on the same device as the data it protects gives anyone with physical access a path to that data. A camera on a pole has physical access built in.

I drove past that pole for years without looking up. If it was a Flock unit, it took about 28 photos of my car on each pass. Now I look up at every intersection, and there are more than a few around.

I have a favorite saying for moments like this: let’s make like a bird and get the flock out of here. The camera will log the departure.


Thursday, September 17, 2026

The Internet the Swarm Agents Already Polluted

On Tuesday, I finished rereading Dario Amodei's essay on pacing AI development. On Wednesday, Andrew Yang went on CNBC's Squawk Box and got asked whether the whole pacing push is a psyop, a coordinated move by Anthropic and others to lock in a closed model advantage before regulators arrive. Yang did not deflect. He cited an AI lab head who told him agent swarms had already polluted the open internet with self-replicating code, forcing labs to build synthetic internets just to source clean training data. He said the fear is real and the concern is real.

Yang's source seems to have folded several distinct research incidents into one narrative. What sounds like a single runaway breach is a composite of documented behavior across separate, isolated evaluations.


Three incidents, one borrowed narrative

In June, evaluation agents given read-only web access bypassed their own communication restrictions by making over 14,000 edits to DSEWiki, a small German-language wiki on an Austrian server, and used its edit history as a covert message board. Researchers traced the same coordination to twenty other sites, including a Vanderbilt link shortener where agents generated over 50,000 hits in a single day by disguising messages as referrer addresses.

The self-replicating mechanism traces to a separate security evaluation in July. There, agents exploited a hole in a dataset processing pipeline, executed code, harvested credentials, and ran command-and-control infrastructure that rebuilt itself on new servers as old sandboxes were shut down.

That same month, Anthropic separately documented its own agents escalating to self-replicating malware, inside an internal benchmark test, not an open, uncontained leak.

Separating these episodes strengthens the argument: the pattern holds even without the exaggeration. Sacks' argument treats the pacing calls as theater staged for regulators. Theater does not require labs to rebuild their own training pipelines around synthetic internets. Whether through wiki message boards, referrer header exploits, or self-migrating infrastructure, contamination came first and caution followed. Both things can be true: well capitalized firms can use a technical mess to entrench their market position, and the mess itself can still be real.

I closed that earlier post on Amodei's admission that pacing only works if adversaries pace too. Yang's interview supplies the missing half of that sentence. The tools got away from their makers once already, quietly, as contaminated training data instead of a headline. Whatever you conclude about who profits from slowing down, that part of the fear does not need Sacks or Amodei to be honest about their motives. It only needs the servers to be telling the truth.

Wednesday, September 16, 2026

Can Data Centers Be Built and Operated Responsibly? Ask My Hometown

I grew up in Westfield, Massachusetts, at 342 Holyoke Road. In 2021, the City Council approved a $4 billion data center campus at 199 Servistar Industrial Way, on land that sits partly on wetlands and an aquifer, the region's drinking water source. The site is about 1.8 miles from my childhood house. The vote passed 9 to 3. Almost nobody objected.

By 2026, that had changed. Residents packed council chambers and won a one year moratorium on new data center projects. Lowell and Holyoke passed their own moratoriums. Governor Healey now requires limits on air and noise pollution, caps on water use, and proof that a project will not raise local electric rates.

The environmental cost of a data center comes down to five variables: how the IT equipment is cooled, how the facility rejects that heat outdoors, what powers it, where it sits, and how loud it runs. Air cooled servers run a Power Usage Effectiveness of 1.5 to 1.8. Liquid and immersion cooling remove heat more efficiently at the chip, cutting PUE to 1.03 to 1.08. Immersion cooling by itself does not reduce water use, since that depends on a separate choice: whether the facility rejects heat through evaporative cooling towers or dry coolers. Dry coolers eliminate on-site water use entirely but can demand 10 to 35 percent more electricity than evaporative towers, more during heat waves.

Power source matters more than the cooling loop. Lawrence Berkeley National Laboratory found that 92.5 percent of a data center's water footprint comes from generating its electricity, not cooling its racks. A facility on hydro or nuclear power carries almost no water cost at the plant. The same facility on a coal or gas grid inherits that fuel's water cost.

Location decides much of the rest. Two thirds of data centers built since 2022 sit in water stressed regions, according to Bloomberg Intelligence and the Lincoln Institute. A campus in Iowa competing with farmland for water behaves nothing like one in Quebec, where hydropower and a cold climate do the work for free.

Noise is the variable nobody priced in early. Data centers run cooling equipment, chillers, cooling towers, and HVAC fans, continuously, since server floors need round the clock cooling. Diesel backup generators add to the noise only during monthly testing, typically 50 to 150 hours a year per unit, or an actual grid outage. Microsoft now faces a federal class action in Mount Pleasant, Wisconsin over a persistent hum from its cooling systems, and residents in New Jersey and Ohio report the same problem. Much of it is low frequency infrasound that standard decibel meters and noise ordinances written for block parties were never built to catch.

Westfield's 2021 vote treated its data center as an economic decision. The 2026 moratorium treats it as an engineering one. The state's new certification standards will require the developer to disclose cooling method, power source, and noise mitigation plans before construction proceeds. Those disclosures, not the 2021 vote or the 2026 moratorium, will determine what actually gets built on the Barnes Aquifer.