• 7 mins read
  • Published

Phala builds confidential AI to keep agent data and execution private

Guido Molinari Blockchain economics and tokenomics writer EgonCoin

Post by Guido Molinari

Phala builds confidential AI to keep agent data and execution private EgonCoin © egoncoin.com
Phala builds confidential AI to keep agent data and execution private © egoncoin.com

Phala is rolling out confidential AI tools that keep agent prompts, memory, and API keys hidden from cloud providers and host machines. The system uses hardware-level trusted execution and attestation to cut down on data leaks.

AI agents have moved far beyond simple chatbots. Now they run code, tap into databases, and even handle digital assets. With that power comes new privacy risks. Prompts, memory, API keys, and wallet permissions can all leak if not protected. Phala is betting on confidential AI, using trusted execution environments (TEEs) and GPU hardware isolation to set new rules for trust in cloud AI.

Hardware-level protection

Standard cloud security like HTTPS and database encryption only go so far. Once sensitive data hits a server's memory for AI inference, users have to trust the cloud provider's setup, hypervisor, and admins to keep it private. Confidential AI tries to shrink that trust zone by running data inside hardware-protected spaces. TEEs carve out isolated areas at the CPU or GPU level. They encrypt memory and check integrity, so even privileged software on the host can't see what's running inside.

Phala's TEE-backed AI infrastructure processed over 202 billion tokens in a single 24-hour period, highlighting real-world demand for privacy-first AI services.

Analyst

Big AI models need GPU TEEs. NVIDIA's Hopper and Blackwell GPUs now support confidential computing, letting AI inference run inside a hardware-rooted trusted space. Phala's setup combines CPU and GPU TEEs with attestation. This lets users check that their agent's code is running in the right environment before any keys or permissions are handed over. NVIDIA says confidential computing for AI inference is already live on Blackwell GPUs. Confidential VMs and encrypted NVLink keep both data and models safe during processing. This is a big step for protecting enterprise data and intellectual property as AI agents take on more work.

Agent security and attestation

AI agents often need to call outside APIs, use wallets, or store long-term memory. Each step can leak sensitive info if not locked down. Phala puts agents inside attested confidential virtual machines, tying the agent's runtime to its permission scope. Attestation creates cryptographic proofs linked to the hardware and software stack. This lets outside verifiers check that the agent is running the right code in a real TEE before any secrets are shared. In Phala's privacy-first AI inference gateway, every request runs inside a hardware-attested confidential VM. The attestation quote is signed by trusted hardware and tied to the exact container image with a hash.

This setup matters most for Web3 agents, where wallet permissions and signing keys need tight control. By keeping credentials and memory inside a hardware-protected sandbox, Phala aims to block both insider threats and infrastructure bugs from leaking agent data. The attestation process makes sure that if the code or environment changes, sensitive credentials stay locked. This matches industry best practices, where confidential AI is used to release secrets and access only after a successful attestation-not just by flipping on a "confidential mode."

NVIDIA's confidential computing is now positioned as the third generation of confidentiality across Hopper, Blackwell, and Rubin platforms, as highlighted in the VAST DataEnclave announcement. This evolution enables trusted infrastructure for uniting leading AI models and sensitive enterprise data.

VAST Data Press Release

Modern cloud platforms offer encryption, access controls, and network isolation, but users still have to trust the provider's infrastructure once data is in use. Confidential AI doesn't replace cloud computing. Instead, it narrows the trust zone by making sure sensitive work runs in spaces the provider can't access. This doesn't fix every AI security issue-prompt injection, malicious tools, and app-level bugs are still a problem-but it does control who can see and manage the computation.

Performance is still a real concern. Phala's 2026 research found that with Blackwell GPU confidential computing, the main bottleneck for large language models wasn't the GPU, but the data bridge between the confidential VM and GPU. NVIDIA's own tests with DeepSeek-R1 on eight B200 GPUs showed confidential computing throughput at 96.1% to 98.2% of the baseline. Per-token latency went up by 1.2% to 4.3%, depending on concurrency. These numbers show confidential AI is moving from proof-of-concept to production, but there are still trade-offs between security and speed.

Tool calls and verifiable execution

Modern AI agents stand out for tool calling-the ability to trigger search, database, code, or blockchain tools as needed. Confidential AI doesn't block these features. Instead, it brings them into a permission-controlled, attested setup. For example, an agent can read a protected API key inside a confidential VM and call an outside service through a controlled network path, exposing only the needed request, not the agent's full memory or prompt.

Phala's design uses per-agent sandboxes, sealed vaults for credentials, and scoped outbound channels. The aim is to keep agent identity, code, credentials, and tool permissions inside a verifiable execution space. This isn't about hiding chat content. It's about making sure every part of agent operation-from memory to tool calls-is protected and can be checked.

Salesforce's move into AI agents and unified data clouds has shown how complex and risky agent-based automation can get, as reported earlier. Phala's confidential AI model takes a different path, focusing on hardware isolation and cryptographic attestation instead of just platform-level controls.

Confidential AI and standard AI cloud services can work side by side. Both use TLS for data in transit and encryption for storage. Confidential AI adds hardware protection for data in use and lowers the trust needed in the cloud provider. Attestation lets users check the runtime environment, not just take the provider's word for it. Still, confidential AI doesn't automatically stop prompt injection, model hallucinations, or bad permissions-TEEs control where computation happens and who can access it, not whether the model's output is correct.

NVIDIA has rolled out confidential computing to its Hopper, Blackwell, and Rubin GPUs. Apple Private Cloud Compute has also announced using NVIDIA confidential GPUs for server-side inference. These moves show confidential AI is becoming part of the main infrastructure for AI workloads, not just a niche for blockchain or highly sensitive data.

As confidential AI infrastructure grows up, the main challenge is balancing security, speed, cost, and developer experience. Phala's roadmap is to make TEE and GPU confidential computing easy for developers building agents with higher permissions, longer memory, and more complex tool use. The real value will depend on whether these protections can fit into real AI products without slowing them down or making them hard to use.

NVIDIA's September 2026 performance tests showed confidential computing on Blackwell GPUs hit up to 98.2% of baseline output-token throughput for large language model inference. Extra per-token latency ranged from 1.2% to 4.3%, depending on concurrency. These results show hardware-level protection is close to production-ready for heavy AI workloads.

Trusted execution environments (TEEs) are hardware security tech that keeps code and data separate from the rest of the system-even from admins or privileged software. In confidential AI, TEEs protect not just user data but also model weights, prompts, memory, and API keys during runtime. Attestation is the process of creating cryptographic proofs that a workload is running in a real TEE, so outside parties can check the environment before releasing sensitive credentials or permissions. While TEEs can cut the risk of data leaks from infrastructure hacks, they don't fix all app-level bugs. Developers still need to set strong permission controls and check inputs carefully.

Related articles