Cohere's Model Vault Now Encrypts AI Inference So Even Cohere Can't See It

Cohere adds confidential computing to Model Vault, sealing enterprise AI inference inside NVIDIA-Intel-AMD trusted execution enclaves so neither Cohere nor cloud operators can read prompts, embeddings or outputs.

Cohere's Model Vault Now Encrypts AI Inference So Even Cohere Can't See It

Cohere has switched on confidential computing across its Model Vault inference platform, sealing enterprise AI workloads inside hardware-attested trusted execution environments so that even Cohere's own operators cannot read prompts, embeddings or outputs.

Encrypted at rest, in transit — and now in use

Announced September 16, 2026, the Model Vault Encrypted tier runs Command foundation models on confidential CPUs from Intel (TDX) and AMD (SEV-SNP) paired with NVIDIA GPUs in confidential mode. Every inference call returns a cryptographic attestation report that lets customers verify the exact hardware, firmware and security policies protecting their workload before any data is sent.

"Every inference returns an attestation report that lets customers verify the exact hardware, software and security policies protecting their workload," Manoj Govindassamy, Cohere's director of serving inference, said in the launch. Data is decrypted only inside the enclave, and Cohere says it holds no keys or copies capable of accessing customer prompts or completions.

Zero-trust inference for regulated industries

Model Vault is Cohere's dedicated inference platform for enterprises deploying Command, Aya and Rerank models in private, on-premise or sovereign environments. The Encrypted tier is aimed at Cohere's target verticals — banking, insurance, healthcare, manufacturing, government — where data-loss prevention rules and regulator scrutiny have kept sensitive prompts off shared cloud endpoints.

Cohere Model Vault confidential inference platform

An arms race for private inference

Confidential computing has been a background feature in cloud AI for a year, but Cohere is one of the first foundation-model labs to build a fully attested inference product on top of it, offering customers a mathematical guarantee that their AI prompts stay opaque to everyone including the model vendor. Rivals from Anthropic to OpenAI to hyperscalers like Azure and Google Cloud have all telegraphed similar private-compute options, but Cohere is packaging it as the default for its regulated customers.

The launch strengthens Cohere's positioning as the sovereign-and-regulated AI specialist — a wedge the Toronto-Berlin company has sharpened after its Aleph Alpha combination — and follows a wave of dedicated NVIDIA confidential-computing offerings across the AI stack.

Reporting based on coverage from VentureBeat, Cohere and Cohere's technical documentation.

Category: Cyber Security

Tags: Edge Computing AI Cybersecurity Nvidia AI Security

Related Articles