Air-gapped AI security architecture showing thermal covert-channel containment in a modern datacenter

OpenAI’s Air-Gap FUD Fails the Basic Standards of Security Engineering

Research Summary

In a September 17, 2026 interview with Dwarkesh Patel, Noam Brown, a researcher at OpenAI, said he was not convinced that even air-gapping computers would be sufficient, then pointed to research where nearby systems communicated through processor heat and temperature sensors.[1] Brown’s thermal communication example takes a legitimate covert-channel experiment and stretches it far beyond what the research established.

The research he was describing is almost certainly BitWhisper. In 2015, researchers demonstrated thermal communication between two already-compromised computers placed within 40 centimeters of each other. The effective throughput was 1 to 8 bits per hour.[2] Those details completely change the threat model.

Modern AI systems execute across containers, virtual machines, shared accelerators, schedulers, and geographically distributed infrastructure. GPU telemetry can also be restricted. NVIDIA vGPU guests, for example, cannot query the physical GPU’s temperature or power consumption through NVML.[3] Thermal covert channels are real and worth studying, but using BitWhisper to casually suggest that even an air gap may not contain advanced AI is a much larger claim than the research supports.

Introduction

Air-gapped AI security should be judged by the same standards we apply to every other security control. If I walked into a client’s boardroom and told the CEO, “Attackers can bypass your firewall, so I am not convinced firewalls are sufficient,” I would expect to get challenged immediately. The statement is technically true, but technically true is nowhere near enough to support a security finding.

Attackers bypass firewalls through vulnerable permitted services, stolen credentials, trusted endpoints, tunneling, supply-chain compromise, configuration mistakes, and occasionally vulnerabilities in the firewall itself. None of that makes the control meaningless. A pentester still has to establish whether a realistic attack path exists in the environment, what prerequisites it requires, which compensating controls interfere with it, and what impact can actually be demonstrated.

That same standard should apply when OpenAI discusses AI containment. During his interview with Dwarkesh Patel, Noam Brown questioned whether even air-gapping computers would be sufficient because nearby systems can theoretically communicate through changes in temperature.[1] The underlying research is legitimate. The problem is the enormous amount of engineering that disappears between that research and the implication being presented.

For people who do not work with containers, virtualization, GPUs, datacenter cooling, and security architecture, the takeaway is easy to predict. Advanced AI might simply communicate through heat and defeat the air gap anyway. That is exactly why this kind of statement deserves more care.

Technical Deep-Dive

Air-gapped AI security needs an attack path

Brown’s actual statement matters. During the discussion of AI safeguards, he said:

“You could even go as far as to say, ‘Well, we should air gap the computers.’ I’m not convinced that that would be sufficient.”

He then described research where one computer changed its processor temperature and another nearby computer detected those changes through temperature sensors. Brown also acknowledged that the work was “mostly academic.”[1]

The problem is not that he brought up an obscure covert channel. Security researchers should explore obscure covert channels. The problem is using that example to cast doubt on a major containment control without explaining the conditions that made the attack work.

Nobody is claiming an OpenAI model has communicated through heat, and nobody is claiming BitWhisper has been observed as an autonomous AI escape mechanism. That was never the issue. The issue is that the public warning skips directly from “researchers demonstrated this once” to “I am not convinced an air gap would be sufficient.”

That would never be acceptable reasoning in a penetration test. A finding needs an attack path, evidence, context, and some consideration of the compensating controls before risk is assigned. “Technically possible” starts that analysis. It does not finish it.

What BitWhisper actually demonstrated

BitWhisper was published in 2015 by researchers at Ben-Gurion University. The researchers created a bidirectional thermal covert channel between two computers that were already compromised.[2]

One machine varied its computational workload to change its heat output. The receiving computer monitored temperature sensors and decoded those changes. The machines had to be close enough for the heat generated by one to measurably influence the other.

The published operating characteristics were:

Maximum distance: approximately 40 centimeters

Effective throughput: approximately 1 to 8 bits per hour

Sender state: already compromised

Receiver state: already compromised and intentionally monitoring temperature

Communication medium: physical heat transfer between nearby systems

Those are not minor implementation details. They define the attack. BitWhisper demonstrated that an air gap does not eliminate every conceivable physical side channel, which is interesting and useful research.

It did not demonstrate that arbitrary isolated workloads can communicate thermally across modern datacenter infrastructure. It also did not demonstrate that a model can somehow create the missing receiver, telemetry access, physical proximity, or useful bandwidth through intelligence alone.

Put 1 to 8 bits per hour into perspective

The throughput matters because risk depends on what can actually cross the boundary. Using BitWhisper’s reported effective rate gives approximate transfer times like these.[2]

Information communicatedSizeAt 8 bits/hourAt 1 bit/hour
One yes/no decision1 bit7.5 minutes1 hour
One byte8 bits1 hour8 hours
16-byte value128 bits16 hours128 hours
32-byte value256 bits32 hours256 hours
1 KiB8,192 bitsAbout 42.7 daysAbout 341.3 days

A single bit can absolutely matter if the receiver already understands what that bit means. A short value may also be useful in a carefully prepared attack. Once that argument is made, though, the threat model immediately becomes more demanding because something on the other side must already be in a position to receive and use the information.

Who is receiving the signal?

Where is that receiver executing?

Why is it physically close enough to observe the sender?

Which temperature or power measurement can it access?

How does the receiver know when transmission begins?

How long can both workloads remain in place?

How does it distinguish the signal from normal workload changes?

What capability does the receiver already possess?

What does receiving those bits actually allow it to do?

Those questions are not objections around the edges of the attack. They are the attack.

BitWhisper already assumed a cooperating receiver

The receiving side is one of the most important pieces that disappears when BitWhisper is reduced to “computers can communicate through heat.” The original research assumed malicious software on both systems.[2]

Heat did not compromise the receiving machine. The sender did not raise the temperature until a neighboring computer somehow became cooperative. The receiver was already present, already controlled, already reading the appropriate measurements, and already using the agreed encoding.

Translated into an AI containment scenario, the threat profile starts looking considerably more specific:

Required thermal-channel threat profile

1. A sender exists inside the restricted environment.
2. The sender can repeatedly influence a useful physical resource.
3. A cooperating receiver already exists across the intended boundary.
4. The receiver can observe a temperature-related signal.
5. Sender and receiver are physically coupled strongly enough for the signal to survive.
6. Both workloads remain available long enough to exchange useful information.
7. The receiver understands the encoding.
8. The information received gives the receiver some useful capability.

If those conditions exist, the channel deserves investigation. If they do not, pointing to BitWhisper does not make them appear.

There may be architectures where some of those prerequisites are realistic. Two untrusted workloads might share a GPU. Multiple agents could intentionally operate in different security zones. Another host might already be compromised, or a poorly designed telemetry service might expose physical measurements too broadly. Those are actual threat models that can be tested. “Advanced AI will find a way” is not one.

An AI shell is not necessarily running beside the AI’s GPUs

Another problem with the thermal argument is treating “the AI” like a process running on a single physical computer. That is not how modern AI services are generally constructed.

The environment where an agent executes tools may be completely separate from the infrastructure performing model inference. An agent might receive a short-lived CPU container for shell execution while the model itself runs across GPU infrastructure elsewhere. OpenAI’s own shell tooling supports hosted containers as well as developer-controlled local runtimes.[6]

If an agent executes:

python exploit.py

the resulting computation affects the hardware running that tool environment. It does not automatically affect the GPU that happened to process the model’s previous inference request.

This matters because a usable thermal transmitter requires a reasonably stable relationship between logical workload activity and a particular physical resource. Modern inference infrastructure may batch requests from many customers, route jobs among accelerators, change scheduling assignments, or share devices among workloads. The agent may be capable of requesting more computation without knowing which physical GPU will perform it.

A thermal receiver has the inverse problem. It needs a stable physical signal that corresponds closely enough to the sender’s activity to decode useful information. Calling both sides “AI” does not establish that relationship.

Virtualization does not eliminate heat, but it absolutely changes the attack

BitWhisper’s researchers tested heat generation from a VirtualBox VM. The VM received four of the host’s eight logical processors. The researchers observed approximately three minutes of propagation delay and a maximum receiver temperature increase of about 3 degrees Celsius compared with 4 degrees in the native experiment.[2]

That result is useful because it confirms something obvious but important. A virtual machine still performs computation on physical hardware, and containers do the same. Virtualization does not repeal thermodynamics, but it does change resource control, placement, persistence, and visibility.

A container may have a CPU quota. A VM may receive only part of a host’s processing capacity. A GPU scheduler may divide work among tenants. A workload may move, and another tenant may consume resources as soon as one workload becomes idle. The receiver may not remain on the same physical machine long enough to establish a stable baseline.

Docker supports CPU quotas and CPU affinity for exactly this kind of resource governance. Containers have no hard CPU limit by default, but production deployments can constrain them significantly.[7]

The real question therefore becomes much narrower and much more useful:

Can the isolated workload intentionally modulate a stable physical resource
strongly enough, and for long enough, that another isolated workload can
observe and decode the resulting change?

That question can be tested. The statement “AI can create heat” tells us almost nothing.

Sensor access is an enforceable boundary

The receiving side of the attack also needs something to measure. It is tempting to assume that because physical hardware has temperature sensors, software running on the hardware can simply read them. That is not how virtualization boundaries work.

NVIDIA’s Container Toolkit exposes GPU functionality according to the devices and driver capabilities assigned to the container. Its utility capability enables interfaces such as nvidia-smi and NVML.[8] A GPU-enabled container may therefore expose useful telemetry if the environment is configured that way.

That is a configuration decision, and NVIDIA vGPU demonstrates the opposite case. NVIDIA explicitly documents that physical temperature information is not available through the vGPU guest interface. Calls such as:

nvmlDeviceGetTemperature()

return:

NOT_SUPPORTED

The same applies to physical GPU power consumption through calls such as:

nvmlDeviceGetPowerUsage()

[3]

The physical GPU obviously still has temperature and power measurements. The guest simply does not receive them through that interface. That distinction is fundamental to security engineering because a capable model cannot reason its way through a missing API response unless there is some other accessible path.

If another path exists, find it. If a vulnerability exposes the measurement, demonstrate it. If another sensor leaks equivalent information, test that sensor. The existence of a physical sensor does not establish that the isolated workload can observe it.

Test the environment instead of theorizing past it

If thermal communication were considered a meaningful risk during an authorized AI containment assessment, the first task would be to inventory what the workload can actually observe. For a GPU environment, a simple query might begin that process:

nvidia-smi --query-gpu=uuid,temperature.gpu,power.draw --format=csv

The result tells us something concrete about that execution environment. If physical temperature and power are exposed, that becomes part of the attack surface. If the fields return N/A, NOT_SUPPORTED, or no compatible device exists, that particular observation path is unavailable.

The next step would be identifying what the workload actually received:

Telemetry validation profile

Physical GPU or virtual GPU?
Passthrough device or mediated device?
MIG instance?
Stable device assignment or dynamically scheduled?
Temperature visible?
Power visible?
Fan telemetry visible?
Host hardware-monitoring interfaces visible?
Sampling frequency sufficient for meaningful observation?
Same physical allocation retained long enough to establish a signal?

This is how a theoretical concern becomes a security assessment. You test the prerequisites instead of assuming past them.

Now put the experiment inside a datacenter

BitWhisper involved systems separated by tens of centimeters. Modern AI infrastructure sits inside datacenters specifically engineered to control very large amounts of heat, and that difference matters because heat is the communication medium.

Datacenter cooling determines where heated exhaust travels, how quickly it is removed, how much air is recirculated, and whether one server’s exhaust meaningfully affects another server’s intake. Hot aisle containment exists specifically to prevent hot exhaust from mixing freely with cold supply air. Department of Energy case studies describe these techniques as core datacenter efficiency and cooling controls.[11]

BitWhisper itself demonstrated that physical placement affected the signal. In one configuration, the receiver took approximately 25 minutes merely to register a 1 degree Celsius temperature increase.[2]

A simplified relationship for an air-mediated channel looks like this:

Receiver intake temperature rise ≈

    fraction of sender heat reaching receiver
    × additional sender power
    -----------------------------------------
    air mass flow × air heat capacity

The critical value is the fraction of the sender’s additional heat that actually reaches the receiver. Suppose, purely as an illustration, a workload adds 100 watts of heat. If only 1 percent reaches a neighboring receiver, airflow is 0.1 kilograms per second, and the heat capacity of air is approximately 1,000 joules per kilogram-kelvin, the resulting intake change is roughly:

0.01°C

That number is not intended to model a specific datacenter. It demonstrates why the physical environment cannot be ignored.

A rack full of GPUs consuming enormous amounts of power does not automatically create a useful covert channel. The relevant quantity is the additional heat the sender can intentionally manipulate and how much of that variation reaches something the receiver can measure. Those are very different questions.

Liquid cooling changes the physical channel again

High-density AI infrastructure increasingly uses direct liquid cooling. NVIDIA’s rack-scale systems, for example, move significant amounts of heat from GPUs and networking components into coolant rather than relying entirely on server airflow.[12]

That does not make every thermal side channel impossible. It means the proposed attack has changed.

If the theory becomes that one workload alters coolant temperature and another workload observes the difference, the assessment now needs a different set of facts:

Liquid-cooled thermal-channel questions

Which components share a coolant loop?

Where are temperatures measured?

What sampling resolution is available?

Can the isolated workload access those measurements?

How quickly does coolant respond to workload variation?

How much thermal mixing occurs before the measurement point?

Can one workload remain on the same physical equipment long enough?

Can normal rack activity overwhelm the signal?

Those would be interesting research questions, but they are nowhere near the original BitWhisper experiment. Two air-cooled desktop systems sitting a few inches apart do not become a model for liquid-cooled AI racks merely because both systems ultimately produce heat.

Geography makes the blanket claim even weaker

Large AI platforms are distributed systems. Two agents that interact logically do not necessarily execute on the same GPU, server, rack, room, datacenter, or continent. That becomes rather important when the proposed communication medium is heat.

Physical placement changes the threat

Same physical GPU:
A shared-device side channel may be plausible and deserves testing.

Same server:
Shared chassis, sensors, and power behavior may provide observables.

Separate servers in one room:
Heat must survive the real airflow and cooling architecture.

Separate datacenters:
The BitWhisper-style local thermal path no longer connects the endpoints.
A separate relay mechanism would be required.

Once another relay mechanism is introduced, that mechanism becomes part of the attack path and needs to be analyzed on its own. The fact that several workloads belong to the same AI system does not create physical adjacency.

Likewise, increasing the number of agents does not guarantee that one of them happens to occupy a useful receiving location with access to the right measurement. Agent swarms may change the probability of interesting placement conditions in some architectures, but the architecture still has to provide those conditions.

Shared-GPU thermal channels are real and worth separating from BitWhisper

There is newer thermal-channel research that is considerably more relevant to shared accelerator environments. A 2023 DATE paper demonstrated communication between cooperating applications sharing a GPU. The researchers reported maximum rates of 1.38 bits per second on an RTX 2070 and 8.75 bits per second on a Jetson TX2.[4]

That is dramatically faster than BitWhisper, but it is also a very different attack. The sender and receiver shared a physical GPU. Heat did not have to travel between independent computers across a room. The research concerned cross-application leakage through shared hardware.[4]

That is a legitimate cloud-isolation concern. If workloads from different trust zones share a physical accelerator and both can observe thermal behavior with enough fidelity, that is something I would want evaluated.

The appropriate security conclusion would be narrow:

Logical GPU isolation should not automatically be assumed to eliminate
physical side channels between workloads sharing the same device.

That conclusion is useful because it gives defenders something specific to test. It does not require concluding that air-gapped AI systems are broadly uncontainable.

MIG does not magically answer the question either

NVIDIA Multi-Instance GPU, or MIG, partitions supported GPUs into isolated GPU instances with dedicated compute and memory resources.[13] Those instances still reside on the same physical GPU, which makes shared physical effects a reasonable research question.

At the same time, sharing hardware does not automatically expose useful thermal telemetry or prove that a receiver can distinguish another workload’s activity. A meaningful assessment would need to establish the entire path:

Shared-GPU threat validation

1. Do workloads from different trust zones share the same physical GPU?
2. Can the sender influence power or temperature in a repeatable way?
3. Can the receiver observe a related signal?
4. Is the receiver's measurement sufficiently precise?
5. Does normal system activity obscure the signal?
6. How much information survives?
7. Does that information cross a security boundary that matters?
8. What capability does the receiver gain as a result?

That is the difference between security testing and storytelling. The result may be concerning, but if it is, the evidence will show why.

“The AI is really smart” is not an exploit primitive

There is a recurring problem in AI safety discussions where intelligence quietly becomes the answer to missing attack prerequisites. The control blocks access, but perhaps the AI will find a way around it. The sensor is unavailable, but perhaps the AI will discover another one. The systems are physically separated, but perhaps multiple agents will somehow bridge the gap.

A sufficiently capable model may absolutely discover vulnerabilities that humans missed. It may identify strange interactions between systems, exploit implementation mistakes, or combine small weaknesses into attack chains that nobody anticipated. That is precisely why adversarial testing matters.

But intelligence cannot replace the attack chain in the analysis. When a pentester documents an exploit chain, every transition needs some mechanism behind it.

Initial capability
        ↓
Reachable interface
        ↓
Vulnerability or permitted behavior
        ↓
New capability
        ↓
Security boundary crossed
        ↓
Demonstrated impact

The analysis does not remain valid if several transitions are replaced with:

The attacker is extremely capable and might figure something out.

That standard would make every control indefensible because every control is theoretically vulnerable to something not yet discovered. Security engineering cannot work that way.

Compensating controls are the entire point

Brown’s framing also treats the air gap as though it must independently defeat every imaginable attack, which is not how security architecture works. Controls stack, and each control changes what the attacker has to accomplish next.

A firewall is not expected to stop every possible attacker behavior by itself. Authentication does not eliminate every path to unauthorized access. Endpoint protection does not make exploitation impossible, and segmentation does not guarantee that no information can ever move between zones.

A serious AI containment architecture could combine several layers:

Example containment control stack

Physical network separation
Strict workload isolation
Ephemeral execution environments
CPU and GPU resource quotas
Restricted hardware telemetry
Dedicated physical hardware
Controlled workload placement
Cooling and airflow separation
No persistent cooperating receiver
Short execution windows
Behavioral monitoring
Independent model-level controls

An attacker has to defeat the architecture that actually exists. If the thermal attack depends on high CPU utilization but the sender receives a small quota, that matters. If it depends on GPU telemetry that the guest cannot read, that matters. If the receiver is in another facility, that matters. If execution lasts five minutes and the proposed channel needs hours, that matters.

Those controls do not need to make thermal communication impossible in every conceivable environment. They need to reduce the actual risk in the system being protected. This is routine security engineering everywhere else, and AI should not get a special exemption.

NIST already gives us the right framework

Covert channels are not a new problem created by AI. NIST SP 800-53 addresses covert-channel analysis under SC-31.[14] The control calls for identifying covert channels and estimating their bandwidth. SC-31(3) goes further and calls for measuring bandwidth in the operational environment because laboratory results may differ substantially from operational behavior.

That is exactly how the BitWhisper question should be handled:

BitWhisper gives us a hypothesis.

Architecture determines whether the prerequisites exist.

Testing determines whether the channel survives.

Measurement determines the usable bandwidth.

Impact analysis determines whether anyone should care.

The laboratory result does not automatically become the result for a hyperscale AI environment. If OpenAI believes advanced models could use thermal channels to defeat realistic air-gapped containment, that would be extremely interesting research. Build the environment, expose only the capabilities the contained model would actually receive, and see whether it can establish the channel.

That result would tell us something. An anecdote about two computers on a desk does not.

The standard should be higher when the audience cannot evaluate the claim

Most people hearing an OpenAI researcher talk about AI containment do not spend their days working with vGPU boundaries, MIG, container namespaces, workload schedulers, hardware-monitoring interfaces, or datacenter cooling. They are unlikely to know that BitWhisper transmitted 1 to 8 bits per hour, that the computers were already compromised, or that the systems were within 40 centimeters of each other.

They may also not realize that virtualization can remove access to the physical telemetry the attack depends on. What they hear instead is something much simpler:

Even if we physically isolate advanced AI from the network,
it may still communicate through heat.

That sounds extraordinary because it is extraordinary. Coming from an OpenAI researcher, it also carries authority that the original academic experiment does not have with the general public.

That is why the qualifications matter. We should absolutely investigate weird covert channels. We should assume future models may find relationships between systems that humans overlooked, and we should test AI containment more aggressively than ordinary application isolation because the workload itself may actively search for ways out.

What we should not do is treat the existence of an exotic laboratory attack as evidence that conventional controls may be broadly ineffective without first doing the engineering required to connect the two. That is how legitimate security research turns into FUD.

Insights, Recommendations, and Conclusion

The security industry already knows how to deal with technically possible but unusual attacks. We identify the threat model, enumerate prerequisites, evaluate the control stack, reproduce the behavior where appropriate, measure the resulting capability, and then assign risk based on what survives. There is no reason to throw that discipline away because the attacker happens to be an advanced AI system.

For a thermal-channel assessment, the work should begin with architecture rather than speculation. Determine where tool execution actually occurs, which physical resources remain stable enough to manipulate, and whether another workload exists in a useful receiving position. Test telemetry from inside the same isolation boundary the model receives. Confirm whether the workloads share a GPU, chassis, airflow path, cooling loop, or nothing physical at all. Then measure what happens under normal operational noise.

The complete assessment question is straightforward:

Can this specific sender, using only the capabilities actually available to it,
create a physical signal that this specific receiver can observe across the
intended security boundary, at useful bandwidth, for long enough to produce
a meaningful security consequence despite the compensating controls already
in place?

If the answer is yes, that is a fascinating finding and one worth taking very seriously. If the answer depends on two already-compromised computers sitting 40 centimeters apart, deliberately monitoring temperature, and exchanging a few bits per hour, that context belongs right beside the scary part.

That is where I have the problem with Brown’s statement. As a pentester, I could not present a CEO with a theoretical firewall bypass, ignore the environment and compensating controls, and then conclude that I was not convinced the firewall would be sufficient. That finding would get destroyed in technical QA, and rightly so.

OpenAI should hold public claims about AI containment to at least that standard. Thermal covert channels are real, they deserve research, and they may prove relevant in specific shared-hardware designs. Increasingly capable models may also discover implementations that surprise us. None of that makes it responsible to leap from BitWhisper to broad doubt about air-gapped containment without showing the attack path that connects them.

If we are going to tell the public that even an air gap may not contain advanced AI, then show us the architecture, the prerequisites, the compensating controls that fail, and the channel that still works. Until then, BitWhisper is an interesting covert-channel experiment, not evidence that sound isolation architecture has somehow become obsolete.

Key Takeaways

  • Treat BitWhisper accurately. It demonstrated approximately 1 to 8 bits per hour between already-compromised computers placed within about 40 centimeters of each other.
  • Require a complete attack path before treating a thermal covert channel as a meaningful AI containment risk, including the sender, receiver, physical coupling, telemetry access, usable bandwidth, and resulting impact.
  • Validate hardware telemetry from inside the actual container, VM, vGPU, or other isolation boundary rather than assuming that physical sensors are visible to the workload.
  • Evaluate the entire containment architecture, including resource controls, workload placement, hardware sharing, cooling design, execution lifetime, physical separation, and the existence of a cooperating receiver.
  • Hold public AI safety claims to the same evidentiary standard expected of professional security findings. Technical possibility alone is not a risk assessment.

References

[1] Noam Brown – Agent swarms, alignment, & recursive self-improvement, Dwarkesh Patel, September 17, 2026 – https://www.dwarkesh.com/p/noam-brown

[2] BitWhisper: Covert Signaling Channel between Air-Gapped Computers using Thermal Manipulations, Guri et al., 2015 – https://arxiv.org/abs/1503.07919

[3] Managing vGPUs from a guest VM, NVIDIA Virtual GPU Software Documentation – https://docs.nvidia.com/vgpu/19.0/grid-management-sdk-user-guide/manage-vgpus-guest-vm-sdk.html

[4] The First Concept and Real-world Deployment of a GPU-based Thermal Covert Channel: Attack and Countermeasures, DATE 2023 – https://past.date-conference.com/proceedings-archive/2023/DATA/379.pdf

[5] BitWhisper: Putting the Heat on Air-Gapped Computers, Americans for Ben-Gurion University, March 24, 2015 – https://a4bgu.org/bitwhisper-putting-the-heat-on-air-gapped-computers/

[6] Shell, OpenAI API documentation – https://developers.openai.com/api/docs/guides/tools-shell

[7] Resource constraints, Docker Engine documentation – https://docs.docker.com/engine/containers/resource_constraints/

[8] Specialized Configurations with Docker, NVIDIA Container Toolkit documentation – https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/docker-specialized.html

[9] Naming and data format standards for sysfs files, Linux kernel hardware-monitoring documentation – https://docs.kernel.org/hwmon/sysfs-interface.html

[10] NVIDIA System Management Interface documentation – https://docs.nvidia.com/deploy/nvidia-smi/index.html

[11] 2018 Award-Winning Champion Lessons Learned Case Study: U.S. Department of Energy Thomas Jefferson National Accelerator Facility Data Center Optimization – https://www.energy.gov/cmei/femp/2018-award-winning-champion-lessons-learned-case-study-us-department-energy-thomas

[12] Hotter Than a Hot Tub: The 45°C Breakthrough to Cool AI’s Biggest Machines, NVIDIA, June 21, 2026 – https://blogs.nvidia.com/blog/liquid-cooling-ai-factories/

[13] Introduction, NVIDIA Multi-Instance GPU User Guide – https://docs.nvidia.com/datacenter/tesla/mig-user-guide/latest/introduction.html

[14] Security and Privacy Controls for Information Systems and Organizations, NIST SP 800-53 Revision 5, SC-31 and SC-32 – https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-53r5.pdf