11 min read

Part 7 — AI Compute Is Becoming a Stealable Cyber Capability

From cryptojacking to inferencejacking: attackers can turn exposed AI infrastructure into reconnaissance, reasoning and attack capability. For CISOs, AI usage telemetry is becoming security telemetry—and inference itself a cyber asset worth protecting.
Part 7 — AI Compute Is Becoming a Stealable Cyber Capability

DEF CON 34 — Part 7 of 13

When inference infrastructure becomes attacker infrastructure

By Eckhart Mehler for CISOsCISO — a perspective on cybersecurity leadership, governance and the decisions that determine whether organizations retain control.


For years, defenders have understood stolen compute through a relatively simple economic model.

An attacker compromises infrastructure.
The attacker consumes CPU cycles.
Cryptocurrency is mined.
The victim pays the electricity and cloud bill.

The resource being stolen was computation.

DEF CON 34 suggests that we now need to think about a more consequential version of the same problem.

Attackers can steal inference.

And inference is not merely compute.

It is a capability.

A compromised inference endpoint can provide an attacker with access to systems capable of analyzing targets, generating reconnaissance, interpreting results, planning attacks, assisting exploitation, processing stolen information and supporting lateral movement.

The victim may provide the infrastructure.

The victim may provide the credentials.

The victim may even provide the AI model.

But the attacker decides what the intelligence is used for.

That creates an emerging form of cyber resource theft:

AI compute can become attacker infrastructure without the attacker owning the AI infrastructure.

This is one of the more strategically important implications of DEF CON 34 because it connects three disciplines that most enterprises still treat separately:

AI Security.
Cloud Security.
FinOps.

The telemetry showing how an organization consumes AI may increasingly tell us something about whether someone else is using that AI against us—or against somebody else.


From Cryptojacking to Inferencejacking

The analogy with cryptomining is useful, but only up to a point.

Cryptojacking converts stolen infrastructure into money.

Inference hijacking converts stolen infrastructure into operational capability.

The distinction matters.

Consider the two attack models.

Traditional resource hijacking

Compromised Cloud Account
        ↓
Compute Instance
        ↓
CPU / GPU Consumption
        ↓
Cryptocurrency Mining
        ↓
Attacker Profit

The attacker’s objective is primarily economic.

Now consider stolen inference.

Compromised Credential
        ↓
Inference Endpoint
        ↓
LLM / Agent Infrastructure
        ↓
Reconnaissance
        ↓
Attack Planning
        ↓
Post-Exploitation Analysis
        ↓
Lateral Movement

The resource being consumed is no longer simply computational capacity.

It is machine reasoning capacity embedded inside an attack workflow.

That changes the security significance of unauthorized AI consumption.

A spike in token usage might look like a FinOps problem.

It may actually be an intrusion signal.


Living Off Someone Else’s Inference

One of the most interesting demonstrations around DEF CON 34 captured this idea almost perfectly.

Researchers Redon Gashi and Armend Gashi presented “Living Off Someone Else’s Inference.”

Their premise was straightforward.

Attackers increasingly use LLM-powered tools for activities ranging from reconnaissance automation to post-exploitation enumeration.

But using commercial AI infrastructure creates two disadvantages for attackers:

cost and attribution.

Someone must create an account.

Someone must supply credentials.

Someone must pay for tokens.

Those relationships potentially create evidence.

So why use your own inference infrastructure if somebody else’s is available?

The researchers focused on exposed and recoverable AI resources including:

  • misconfigured Ollama instances,
  • exposed vLLM deployments,
  • leaked AI API keys,
  • credentials discoverable through exposed resources,
  • and authentication material associated with AI coding environments.

Their research introduced two tools: Infreerence, designed to discover and validate usable inference resources across multiple providers, and Echidna, a Mythic C2 agent type capable of consuming discovered inference and using it for operator-directed activities including reconnaissance, exploitation planning, post-exploitation and lateral movement. (⁠Recon Village)

That combination deserves attention beyond the individual tools.

It demonstrates an architectural transition.

The attacker no longer necessarily needs:

Attacker Infrastructure
        ↓
Attacker AI Account
        ↓
Attacker API Key
        ↓
Attacker-Funded Inference

The attack model becomes:

Victim Credential
       or
Exposed Endpoint
        ↓
Victim / Third-Party AI Infrastructure
        ↓
Inference
        ↓
Attacker Workflow

This is the AI equivalent of living off the land.

Except the land now thinks.


Inference Is an Attack Resource

Security teams should resist reducing this problem to stolen API keys.

The credential is only the access mechanism.

The strategically relevant asset is the capability behind it.

An API key might provide access to millions of tokens of inference.

An exposed endpoint might provide access to GPUs.

A compromised AI platform might expose multiple models.

A compromised agent environment might additionally provide tools, memory, connectors and enterprise data.

The attacker therefore acquires something closer to a temporary cyber capability.

That capability can potentially perform:

Target Discovery
        ↓
Data Interpretation
        ↓
Hypothesis Generation
        ↓
Attack Planning
        ↓
Tool Selection
        ↓
Result Analysis
        ↓
Next-Step Recommendation

The attacker still directs the operation.

But increasingly, machine inference can perform some of the analytical work between individual actions.

This is precisely why the DEF CON AI Village’s autonomous-agent experimentation is relevant to enterprise defenders. The Village explicitly provided hosted models and dedicated GPU resources for participants building autonomous pentesting agents. (⁠AI Village)

The important lesson is not that a conference hosted offensive AI experimentation.

It is that the economics are becoming obvious:

inference is useful offensive infrastructure.

Once that is true, exposed inference becomes something attackers have an economic incentive to discover.


The New AI Shadow Infrastructure

Organizations are rapidly accumulating AI infrastructure.

Some of it is obvious:

  • enterprise AI platforms,
  • managed model endpoints,
  • internal copilots,
  • AI gateways,
  • GPU clusters.

Some of it is much less visible:

  • development inference servers,
  • experimental Ollama installations,
  • temporary vLLM deployments,
  • notebooks,
  • model-serving containers,
  • coding-assistant credentials,
  • test API keys,
  • departmental AI projects,
  • abandoned proof-of-concepts.

This creates a familiar cybersecurity phenomenon.

Shadow IT becomes shadow AI infrastructure.

But the risk profile is different.

A forgotten web server exposes an application.

A forgotten inference server exposes computational intelligence.

And because experimentation is currently moving faster than enterprise governance in many organizations, those environments may have weaker authentication, monitoring and network controls than conventional production systems.

The attack surface therefore expands faster than the official AI inventory.

That should concern CISOs.


The Economics Favor the Attacker

There is another reason this attack class matters.

AI-assisted offensive operations consume resources.

Large-scale reconnaissance costs tokens.

Analyzing thousands of files costs tokens.

Generating and evaluating hypotheses costs tokens.

Running multiple agents costs tokens.

Processing post-exploitation data costs tokens.

Attackers therefore face an AI infrastructure bill just as legitimate organizations do.

If they can transfer that cost to somebody else, their economics improve.

Consider an attacker operating across 1,000 targets.

Instead of maintaining:

AI Accounts
+
API Keys
+
GPU Infrastructure
+
Model Hosting
+
Token Budget

they discover usable inference capacity across compromised organizations.

The attacker effectively creates a distributed inference pool.

Victim A ──┐
Victim B ──┤
Victim C ──┤
Victim D ──┼──► Attacker AI Workload
Victim E ──┤
Victim F ──┘

This creates an uncomfortable possibility.

Organizations may unknowingly subsidize the automation of attacks.

Potentially even attacks against other organizations.


AI Infrastructure Could Become Part of C2

Echidna makes another development particularly interesting.

Inference can become integrated with command-and-control architecture.

Traditionally, C2 infrastructure coordinates communication between an operator and compromised systems.

AI introduces another component.

Operator
   │
   ▼
C2
   │
   ▼
Agent
   │
   ├────► Tools
   │
   └────► Inference
             │
             ▼
      Analysis / Planning
             │
             ▼
          Action

Inference does not replace C2.

It augments it.

The model can interpret information collected from the environment and help determine subsequent actions.

This matters because defensive architectures generally assume that enterprise AI services serve enterprise users.

That assumption may become unsafe.

A perfectly legitimate inference request can be malicious in purpose.

The infrastructure sees:

Valid API Call
Valid Credential
Supported Model
Normal HTTPS
Successful Response

The security question is different:

Who is benefiting from the reasoning being performed?

That is significantly harder to determine.


The Attribution Problem

Inference hijacking also creates an interesting attribution problem.

Imagine an enterprise AI endpoint producing requests that appear to involve:

network enumeration,

credential analysis,

Active Directory structures,

vulnerability research,

PowerShell generation,

cloud privilege relationships,

or lateral movement planning.

Whose activity is it?

A penetration tester?

A developer?

A security researcher?

An internal security agent?

A compromised employee account?

An external attacker using a leaked API key?

A malware implant invoking enterprise inference?

The model provider may only see the enterprise tenant.

The enterprise may only see legitimate authentication.

The attacker disappears behind someone else’s identity and someone else’s AI infrastructure.

This creates what we might call:

Inference Laundering

The attacker externalizes three things simultaneously:

Compute.
Cost.
Attribution.

That makes stolen inference particularly attractive.


AI FinOps Is Becoming Security Telemetry

This is where the CISO implication becomes especially interesting.

Organizations building AI platforms already collect extensive consumption telemetry for financial reasons.

They monitor:

tokens,

requests,

models,

GPU utilization,

latency,

cost per application,

cost per department,

and capacity consumption.

Historically, these are FinOps metrics.

Increasingly, they are also security signals.

Suppose a development team normally consumes:

2 million tokens/day

and suddenly consumes:

47 million tokens/day

The first reaction may be:

What happened to the budget?

The security question should be:

Who is using our inference?

The anomaly might indicate:

  • leaked API credentials,
  • compromised service identities,
  • exposed inference endpoints,
  • runaway agents,
  • unauthorized internal applications,
  • compromised coding assistants,
  • malicious automation,
  • or external use of enterprise AI infrastructure.

The important conceptual shift is:

Unexpected AI consumption should be investigated like unexpected cloud execution.

Organizations learned this lesson with cryptomining.

Unexpected EC2 activity became a security signal.

Unexpected GPU and token consumption should increasingly receive similar treatment.

Identity
   +
Inference Endpoint
   +
Model
   +
Token Consumption
   +
Source Network
   +
Application
   +
Tool Calls
   +
Data Access
   +
Time

This allows defenders to ask better questions.

Why is a CI/CD service account suddenly calling a reasoning model?

Why is an inference endpoint normally accessed from Europe receiving sustained requests from unfamiliar infrastructure?

Why is a coding assistant credential producing traffic while the developer is offline?

Why has a low-volume internal application suddenly started using models associated with complex reasoning?

Why is an AI workload consuming tokens while simultaneously generating unusual cloud or network activity?

No individual signal proves compromise.

The composition might.

And that returns us to the central thesis of this entire DEF CON 34 series.

The attack surface is becoming the system.


AI Security Needs Resource Governance

Most enterprise AI security programs currently focus heavily on information.

Can sensitive data enter the model?

Can prompts leak information?

Can users bypass guardrails?

Can generated content violate policy?

These are legitimate concerns.

But AI systems expose another dimension.

They provide resources.

An AI platform therefore needs security controls comparable to other high-value infrastructure.

CISOs should increasingly ask:

Who may consume inference?

Not merely who can log into the AI application.

Which humans, agents, workloads and service identities can invoke models?

From where?

Are inference APIs internet accessible?

Are network boundaries enforced?

Can workloads call them directly?

Through which identity?

User identity?

Application identity?

API key?

Managed identity?

Workload identity?

How much?

Are quotas defined?

Are per-identity limits available?

Are abnormal increases detected?

For what workload?

Can inference calls be associated with the application or agent that generated them?

With which downstream authority?

Does the AI merely answer questions?

Or can it invoke tools?

The last question changes everything.


The Capability Multiplier

The risk increases sharply when stolen inference is connected to tools.

An exposed language model is useful.

An exposed agent can be substantially more useful.

Consider the difference.

Model
  ↓
Produces Information

versus:

Agent
  ↓
Inference
  ↓
Tool Selection
  ↓
API
  ↓
Cloud / Network / Repository

Now combine this with Part 3 of this series.

An agent may inherit transitive authority through MCP, cloud identities and connected services.

And combine it with Part 6.

Automation may continuously generate fresh credentials.

The resulting attack path might look like:

Leaked AI Credential
        ↓
Enterprise Agent
        ↓
Inference
        ↓
MCP Tool
        ↓
Workload Identity
        ↓
Cloud Resource
        ↓
New Credential
        ↓
Lateral Movement

No individual component needs to be catastrophically vulnerable.

The attacker exploits the composition.


The CISO Control Model

Organizations deploying substantial AI infrastructure should treat inference capacity as a governed cyber resource.

At minimum, I would introduce six controls.

1. Build an Inference Asset Inventory

Know every environment capable of serving models.

Include:

managed AI platforms,

GPU clusters,

Ollama,

vLLM,

development endpoints,

test environments,

coding assistants,

and departmental deployments.

If it can provide inference, it belongs in the security inventory.


2. Eliminate Anonymous Inference

Production inference should require strongly attributable identity.

Avoid long-lived static API keys wherever stronger workload or federated identity mechanisms are available.

Credential lifecycle management should include AI credentials explicitly.


3. Establish Inference Baselines

Measure normal consumption by:

user,

application,

agent,

service identity,

model,

environment,

and business function.

Without a baseline, abnormal consumption becomes difficult to distinguish from rapid AI adoption.


4. Connect FinOps to the SOC

AI consumption anomalies should be available to security monitoring.

That includes:

unexpected token spikes,

GPU utilization anomalies,

new model consumption,

new source locations,

unusual service identities,

and abnormal invocation times.

FinOps should not discover inference theft three weeks later on an invoice.


5. Correlate Inference With Actions

The most valuable detection will connect:

Inference
     ↓
Agent
     ↓
Tool
     ↓
Identity
     ↓
Action

If a burst of reasoning requests is immediately followed by unusual cloud enumeration or repository activity, the relationship matters more than either event individually.


6. Introduce AI Kill Switches

Organizations need the ability to rapidly revoke:

AI API credentials,

agent identities,

model endpoint access,

tool connectivity,

MCP access,

and inference routes.

A compromised AI environment should not require shutting down the entire enterprise AI platform.

Segmentation matters.


A New SOC Question

Security operations teams will therefore need to add a question to incident investigation.

They already ask:

Was data stolen?

Were credentials stolen?

Was compute compromised?

Increasingly they should also ask:

Was our AI capability stolen?

That question extends beyond financial loss.

The incident may involve the organization’s infrastructure being used to support someone else’s offensive operation.

This creates potential implications for:

incident response,

forensics,

cloud security,

AI governance,

legal teams,

provider relationships,

and threat intelligence.

The logs needed to answer that question must exist before the incident occurs.


The Strategic Shift

There is a tendency to describe enterprise AI primarily as software.

That description is becoming incomplete.

AI infrastructure is simultaneously:

software,
data processing infrastructure,
computational infrastructure,
and decision-support capability.

Agents add another dimension:

operational authority.

Attackers will eventually value each of these.

The evolution therefore looks something like this:

STEAL DATA
     ↓
STEAL CREDENTIALS
     ↓
STEAL COMPUTE
     ↓
STEAL AUTHORITY
     ↓
STEAL INFERENCE

The last two categories are particularly important because they change what attackers can do rather than merely what they can obtain.


The New Economics of Compromise

Cryptojacking taught defenders that cloud infrastructure has value even when no corporate information is stolen.

Inference hijacking extends that lesson.

An AI environment can be valuable to an attacker even if:

no database is exfiltrated,

no ransomware is deployed,

no employee is impersonated,

and no production application is disrupted.

The attacker may simply consume the organization’s reasoning infrastructure.

Quietly.

Perhaps for weeks.

The victim sees an AI bill.

The attacker sees a capability.

That difference in perspective is exactly the type of blind spot attackers exploit.


What CISOs Should Take From DEF CON 34

The important lesson from “Living Off Someone Else’s Inference” is not that exposed Ollama or vLLM instances exist.

Nor is it simply that AI API keys can leak.

Those are implementation problems.

The deeper lesson is architectural.

Once inference becomes operationally valuable to attackers, AI infrastructure becomes part of the resource landscape they will enumerate after compromise.

Attackers already search environments for:

credentials,

cloud roles,

Kubernetes access,

CI/CD secrets,

SSH keys,

storage,

and compute.

We should expect another question to enter that workflow:


But Volume Alone Is Not Enough

Detection cannot rely exclusively on cost anomalies.

A sophisticated attacker does not need to consume enormous amounts of inference.

They may deliberately remain inside normal consumption patterns.

The more interesting telemetry therefore combines several dimensions.

What AI can I use here?

And increasingly:

What can that AI help me do next?

That is a fundamentally different security problem.


The CISO Principle

The transition can be summarized simply.

Yesterday:

Attackers stole CPU cycles to generate money.

Tomorrow:

Attackers can steal inference cycles to generate capability.

That makes AI usage telemetry security telemetry.

It makes AI credentials security credentials.

It makes inference endpoints security assets.

And it makes enterprise AI capacity something that must be protected not only from data leakage or prompt injection, but from unauthorized operational consumption.

Because AI compute is no longer merely a cost center.

It is a stealable cyber capability.


Next in the series: Part 8 — Human Approval Is Becoming Another Attack Surface — why Human-in-the-Loop is not automatically a security control, and how routine approvals can turn human oversight into approval fatigue—making the person meant to be the last line of defense another attack surface.


Publication Note & Disclaimer
This article provides security and governance analysis, not legal advice. Regulatory obligations must be assessed against the facts, jurisdictions, data types, and roles of the organizations involved.

This article reflects my personal professional perspective and does not represent the official policy or position of my employer. Drafting and editorial refinement may have been supported by commercially available AI-assisted tools. The analysis, conclusions and final curation are entirely my own.

For information regarding image credits, copyrights, trademarks and other intellectual property rights, please refer to the Imprint.