CATS Y. Mo
Internet-Draft D. Yang
Intended status: Informational C. Zhou
Expires: 6 April 2027 Huazhong University of Science and Technology
06 October 2026
State, Storage, and Compute Affinity for AI Agent Service
Selection in Computing-Aware Traffic Steering
draft-mo-cats-agent-state-affinity-00
Abstract
AI agent services are stateful and long-running: the speed with which
a step can be served depends on whether the session's context,
retrieved data, tool results, and model-side state such as a
key-value (KV) cache are already available at, or near, the selected
service contact instance, and on whether the computation that the
step needs is ready there. Computing-Aware Traffic Steering (CATS)
exposes computing and network metrics and defines service contact
instance affinity, but it does not expose the availability of
reusable state, does not distinguish a hard locality constraint on
state from a preference for reusing it, does not provide a way to
compare the cost of moving state with the cost of recomputing it, and
does not represent the readiness of a specific computation. This
document describes affinity in three coupled dimensions -- state,
storage, and compute -- for agent service selection. It states the
motivation and the goals of introducing affinity, the requirements
and constraints that affinity places on the mapping of an agent
workload onto storage and compute resources, the metrics and
measurement methods by which affinity is observed, the mechanisms and
the procedure by which affinity is assured, and two cases in detail:
a long-horizon session that passes through several stages, and a
group of similar agents or of tenants served by shared reusable
state. This document defines no wire protocol, no encoding, and no
data model, and it does not define the transfer mechanisms
themselves.
Status of This Memo
This Internet-Draft is submitted in full conformance with the
provisions of BCP 78 and BCP 79.
Internet-Drafts are working documents of the Internet Engineering
Task Force (IETF). Note that other groups may also distribute
working documents as Internet-Drafts. The list of current Internet-
Drafts is at https://datatracker.ietf.org/drafts/current/.
Mo, et al. Expires 3 April 2027 [Page 1]
Internet-Draft Agent Affinity September 2026
Internet-Drafts are draft documents valid for a maximum of six months
and may be updated, replaced, or obsoleted by other documents at any
time. It is inappropriate to use Internet-Drafts as reference
material or to cite them other than as "work in progress."
This Internet-Draft will expire on 3 April 2027.
Copyright Notice
Copyright (c) 2026 IETF Trust and the persons identified as the
document authors. All rights reserved.
This document is subject to BCP 78 and the IETF Trust's Legal
Provisions Relating to IETF Documents (https://trustee.ietf.org/
license-info) in effect on the date of publication of this document.
Please review these documents carefully, as they describe your rights
and restrictions with respect to this document. Code Components
extracted from this document must include Revised BSD License text as
described in Section 4.e of the Trust Legal Provisions and are
provided without warranty as described in the Revised BSD License.
Discussion Venues
Discussion of this document takes place on the CATS Working Group
mailing list (cats@ietf.org), which is archived at
https://mailarchive.ietf.org/arch/browse/cats/.
Table of Contents
Mo, et al. Expires 3 April 2027 [Page 2]
Internet-Draft Agent Affinity September 2026
1. Introduction . . . . . . . . . . . . . . . . . . . . . . . . . 4
2. Conventions and Definitions . . . . . . . . . . . . . . . . . .6
3. Terminology . . . . . . . . . . . . . . . . . . . . . . . . . .6
4. Motivation and Goals . . . . . . . . . . . . . . . . . . . . . 8
4.1. Why Agent State, Storage, and Compute Are Coupled . . . . 8
4.2. Failure Modes without an Affinity View . . . . . . . . . .9
4.3. Goals . . . . . . . . . . . . . . . . . . . . . . . . . .10
4.4. Non-Goals . . . . . . . . . . . . . . . . . . . . . . . .11
5. Problem Statement . . . . . . . . . . . . . . . . . . . . . . 11
5.1. Storage Is Represented Only as Capacity . . . . . . . . .11
5.2. Affinity Is Instance-Level . . . . . . . . . . . . . . . 11
5.3. Transfer and Re-computation Are Not Comparable . . . . . .11
5.4. Compute Affinity Is Not Represented . . . . . . . . . . .11
5.5. Reuse across Similar Agents and Tenants Is Not
Represented . . . . . . . . . . . . . . . . . . . . . . . . 12
5.6. The Long Horizon Amplifies Each of These Gaps . . . . . .12
6. Working Set Elements Relevant to Agent Services . . . . . . . 13
6.1. Element Classes . . . . . . . . . . . . . . . . . . . . .13
6.2. Compute-Side Readiness . . . . . . . . . . . . . . . . . 14
6.3. Why the Properties Are Pairwise . . . . . . . . . . . . .14
7. State, Storage, and Compute Affinity . . . . . . . . . . . . .14
7.1. Affinity, Preference, and Locality . . . . . . . . . . . 14
7.2. How the Three Affinities Interact . . . . . . . . . . . .15
7.3. The Granularity of an Affinity Decision . . . . . . . . .16
8. Requirements and Constraints on the Mapping . . . . . . . . . 17
8.1. The Mapping Relation . . . . . . . . . . . . . . . . . . 17
8.2. Constraint Families . . . . . . . . . . . . . . . . . . .17
8.3. Mapping Requirements . . . . . . . . . . . . . . . . . . 18
9. Affinity Information Requirements . . . . . . . . . . . . . . 20
9.1. State Availability . . . . . . . . . . . . . . . . . . . 20
9.2. Residency Tiers and Retrieval Cost . . . . . . . . . . . 20
9.3. Constraints, Consistency, and Sharing . . . . . . . . . .20
9.4. Compute Affinity and Reuse Scope . . . . . . . . . . . . 21
10. State Handles and Their Relationship to CATS Identifiers . . 21
11. Measuring Affinity: Metrics and Methods . . . . . . . . . . .22
11.1. What Has to Be Measured . . . . . . . . . . . . . . . . 23
11.2. Affinity Metric Catalogue . . . . . . . . . . . . . . . 23
11.3. Metric Semantics and Reporting Standards . . . . . . . .24
11.4. Measurement Methods . . . . . . . . . . . . . . . . . . 25
11.5. Measurement Requirements . . . . . . . . . . . . . . . .26
12. Affinity Assurance: Mechanisms and Procedures . . . . . . . .27
12.1. Mechanism Catalogue . . . . . . . . . . . . . . . . . . 27
12.2. Assurance Procedure . . . . . . . . . . . . . . . . . . 28
12.3. Triggers for Re-evaluation . . . . . . . . . . . . . . .29
12.4. Abort, Failure, and Fallback . . . . . . . . . . . . . .30
12.5. Assurance Requirements . . . . . . . . . . . . . . . . .31
13. Affinity Assurance for Long-Horizon and Multi-Stage
Sessions . . . . . . . . . . . . . . . . . . . . . . . . . . . 32
13.1. Stage Model . . . . . . . . . . . . . . . . . . . . . . 32
13.2. Phase-Differentiated Assurance . . . . . . . . . . . . .32
13.3. Pinning, Retention, and Tier Budget . . . . . . . . . . 33
Mo, et al. Expires 3 April 2027 [Page 3]
Internet-Draft Agent Affinity September 2026
13.4. Checkpoint and Resume after a Pause . . . . . . . . . . 33
13.5. Stage Transitions . . . . . . . . . . . . . . . . . . . 34
13.6. Drift, Decay, and Re-evaluation . . . . . . . . . . . . 34
13.7. Budget Pacing across Stages . . . . . . . . . . . . . . 35
13.8. Failure and Recovery . . . . . . . . . . . . . . . . . .35
13.9. Multi-Agent Stages . . . . . . . . . . . . . . . . . . .36
13.10. Long-Horizon Requirements . . . . . . . . . . . . . . .36
14. Affinity for Similar Agents and for Tenants . . . . . . . . .37
14.1. Similarity and Reuse Groups . . . . . . . . . . . . . . 37
14.2. Elements That Can Be Shared . . . . . . . . . . . . . . 37
14.3. Conditions for Safe Reuse . . . . . . . . . . . . . . . 38
14.4. Tenant Affinity Domains and Isolation . . . . . . . . . 39
14.5. Fairness, Herding, and Quota . . . . . . . . . . . . . .39
14.6. Measurement and Observability across Tenants . . . . . .40
14.7. Requirements for Similar Agents and Tenants . . . . . . 41
15. Interaction with Existing Work . . . . . . . . . . . . . . . 42
15.1. Traceability to the Agent Service Requirements . . . . .43
16. Operational Considerations . . . . . . . . . . . . . . . . . 43
17. Security Considerations . . . . . . . . . . . . . . . . . . .45
18. Privacy Considerations . . . . . . . . . . . . . . . . . . . 45
19. IANA Considerations . . . . . . . . . . . . . . . . . . . . .46
20. Normative References . . . . . . . . . . . . . . . . . . . . 46
21. Informative References . . . . . . . . . . . . . . . . . . . 46
Acknowledgments . . . . . . . . . . . . . . . . . . . . . . . . . 47
Authors' Addresses . . . . . . . . . . . . . . . . . . . . . . . .48
1. Introduction
The CATS framework selects a service contact instance using computing
and network metrics [I-D.ietf-cats-framework]. Agent services add a
consideration that is not represented in that metric set: how much of
what a step needs is already in place at the candidate. A step of an
agent session frequently can be served in two ways at a candidate
instance -- by using state that the instance already holds and by
using the computation that is already ready there, or by obtaining
that state and that readiness again, either by transferring them from
where they are held or by reconstructing them. The difference between
the two paths is often larger than the difference between two
candidate instances on any currently defined metric.
Three dimensions of that consideration are distinguished in this
document. State affinity is the preference for a candidate because it
already holds the working set elements that a step needs. Storage
affinity is the preference for a candidate because the tier at which
those elements are held is usable for the step and because the
durable store that can serve them is close to it. Compute affinity is
the preference for a candidate because the model revision, precision,
adapter, runtime, and accelerator that the step needs are resident
and ready there, so that no artifact load, quantization change, or
capability substitution is required.
Mo, et al. Expires 3 April 2027 [Page 4]
Internet-Draft Agent Affinity September 2026
The three are coupled, and that coupling is the reason this document
treats them together. State can be reused only by a computation that
can consume it, so a state that is present at an instance whose
runtime cannot consume it has no reuse value at that instance.
Changing the instance to gain compute affinity pays for the change in
state: the accumulated state of a session is either moved, which
costs transfer, or rebuilt, which costs accelerator time. Storage
affinity is the bridge between the two, because it determines at
which tier an element is held and how cheaply it can be made usable
by the computation that needs it.
This document states what a CATS system needs to know in order to
take the three affinities into account, and what it has to be able to
do about them. The document is organized around six questions.
Section 4 states the motivation and the goals of introducing
affinity. Section 8 states the requirements and the constraints that
affinity places on the mapping of an agent workload onto storage and
compute resources. Section 11 states how affinity is measured, and
which metric properties make two measurements comparable. Section 12
states the mechanisms by which affinity is assured and the procedure
that applies them. Section 13 applies that procedure to a
long-horizon session that passes through several stages. Section 14
applies it to a group of similar agents and to tenants that share
reusable state.
The gap that motivates the document has been described for one
element of the working set: the KV cache of a large language model,
for which it has been observed that the existing CATS metrics do not
expose cache state and that the distribution framework does not
describe how cached content is distributed or synchronized across
instances [I-D.li-cats-kv-cache-distribution]. The same reasoning
applies to the other elements of an agent working set, and to the
compute side of the same question.
This document assumes the single-domain deployment model of the CATS
framework. The intended standing of the document is informational
groundwork in the sense of the CATS charter [CATS-CHARTER]: it states
what a CATS system has to be able to express, so that the work can be
taken up by the metric definition [I-D.ietf-cats-metric-definition]
and by the data model [I-D.ietf-cats-data-model] rather than by a
protocol extension.
Three distinctions are introduced here that the current documents do
not make.
* A distinction between state availability, which is a quantity, and
state locality, which is a constraint.
Mo, et al. Expires 3 April 2027 [Page 5]
Internet-Draft Agent Affinity September 2026
* A distinction between instance affinity, which is about which
service contact instance serves a session, and state affinity,
which is about where the session's state and its computation are
held.
* A distinction between the cost of obtaining state by transfer and
the cost of obtaining it by re-computation or re-retrieval, and, on
the compute side, between load, which is a property of an
instance, and readiness, which is a property of the pair of a step
requirement and an instance.
2. Conventions and Definitions
The key words "MUST", "MUST NOT", "REQUIRED", "SHALL", "SHALL NOT",
"SHOULD", "SHOULD NOT", "RECOMMENDED", "NOT RECOMMENDED", "MAY", and
"OPTIONAL" in this document are to be interpreted as described in BCP
14 [RFC2119] [RFC8174] when, and only when, they appear in all
capitals, as shown here.
In this document these key words are used to state requirements on
the design of a CATS system that supports agent services. They do not
describe protocol behaviour, and this document defines no protocol,
message format, or data model.
3. Terminology
This document uses the terms of [I-D.ietf-cats-framework] and
[I-D.mo-cats-agent-service-characteristics]. The following additional
terms are used.
State Handle: An identifier under which a working set element can be
addressed. A state handle is not a network address and does not by
itself authorize access to the state that it names.
State Availability: The presence of a working set element at, or
within a reachable tier of, a candidate instance.
Mo, et al. Expires 3 April 2027 [Page 6]
Internet-Draft Agent Affinity September 2026
State Affinity: The preference for a candidate instance because it
already holds, or is able to reach cheaply, the working set elements
that a step needs.
Storage Affinity: The preference for a candidate instance because the
tier at which a working set element is held there is usable for the
step, and because the durable store that can serve the element is
cheap to reach from it.
Compute Affinity: The preference for a candidate instance because the
computation that a step requires is ready there: the model revision,
the precision, the adapter, the runtime, and the accelerator class
that the step requires are resident, so that the step can start
without a load, a conversion, or a substitution.
State Locality Constraint: A rule that forbids a working set element
from being placed, transferred, or reused in a given location or
scope.
Affinity Domain: The set of instances among which a working set
element may be reused, or among which the sessions of a tenant may be
placed, under a stated scope.
Affinity Gain: The reduction in the cost of serving a step that is
attributable to reuse and to readiness, measured against the cost
that the same step would have had if the state and the computation
had been re-established.
Affinity Decay: The reduction of affinity gain over time, caused by
eviction, expiry, replacement of a model or adapter, drift of the
session's requirement, or growth of the working set beyond the
capacity that the holder is willing to commit.
Affinity Budget: The share of the cost, latency, or capacity budget
of a session that the session may spend on establishing, maintaining,
or changing its affinity.
Mo, et al. Expires 3 April 2027 [Page 7]
Internet-Draft Agent Affinity September 2026
Retention Commitment: A statement by a holder that it will keep a
working set element for a stated period or until a stated event,
within a stated capacity limit and subject to stated preemption
rules.
Reuse Group: The set of sessions, steps, agents, or tenants for which
a given working set element may legitimately be reused.
Similar Agent: An agent whose steps require the same computation or
the same reusable elements as another agent, as determined by the
conditions of reuse and not by the identity of the agent.
Stage Transition: The boundary at which the requirement profile of
the next stage becomes known, and at which the placement and the
affinity of the session may be re-established.
4. Motivation and Goals
4.1. Why Agent State, Storage, and Compute Are Coupled
An agent session consumes three kinds of readiness, and each of them
is produced at a cost that the selection decision can either pay or
avoid.
First, the state of the session. A step that can use the accumulated
context of a session avoids re-reading the transcript, re-retrieving
the documents, re-invoking the tools, and re-running the pre-fill. The
alternative to reuse is not a constant: it depends on the length of
the context that has to be re-established, on the throughput of the
candidate at that operation, and on whether the material that has to
be re-retrieved is available at all.
Second, the tier at which that state is held. The same element has
different costs depending on whether it is in accelerator memory, in
host memory, in a node-local store, or only in a shared store that is
reached over the network. The tier determines whether the element is
usable as it stands or whether it has to be copied, paged, or
deserialized before a step can use it, and it determines whether
holding the element competes for the same accelerator capacity that
the step itself needs.
Third, the computation. A step of an agent session is not served by
an arbitrary processor: it needs a model revision, a precision,
sometimes an adapter, a runtime that can consume the state that is
available, and an accelerator class that can execute it. When those
are absent, the step can be served only after an artifact is loaded,
a model is converted, or the step is redirected to a different
capability tier, and each of those is a cost that is paid before the
step begins.
Mo, et al. Expires 3 April 2027 [Page 8]
Internet-Draft Agent Affinity September 2026
The three are coupled because a state is reusable only by a
computation that can consume it, and because a computation that is
ready is worth little without the state to run on. A candidate that
holds the KV state of a session but runs a different model revision
cannot reuse it. A candidate that has the accelerator free but not
the artifact must load it first, and the load costs more than the
difference between many pairs of candidates. This is why an affinity
decision that considers only one of the three dimensions produces a
placement that is worse than one made without considering affinity at
all: it moves the session to gain a dimension and pays for the gain
in the other two.
4.2. Failure Modes without an Affinity View
Six failure modes are observed when affinity is not represented in
the selection input.
* Repeated reconstruction. The same context is prefilled, the same
documents are retrieved, and the same tools are invoked again
because the selection did not know where the results were already
held.
* Migration thrash. Because the cost of moving state is not
comparable with the benefit of the move, sessions are moved for
gains that do not survive the move, and the traffic that the moves
generate degrades the conditions that motivated them.
* Concentration and its correction. If affinity is applied without a
rule that overrides it, the instances that hold popular state absorb
the demand, and if affinity is not applied at all, the state is
re-established everywhere. Both extremes waste capacity in opposite
directions.
* Isolation that is assumed rather than enforced. A reuse key that
is treated as a hint rather than as a scoped identifier can cause
state of one tenant to be reused by another, or a locality rule that
was expressed as a cost to be traded to be violated silently.
* Loss of accumulated work. A long session that pauses and resumes
on an instance that no longer holds its state, or that holds it at
a tier the step cannot use, restarts from a checkpoint that may be
much older than the session itself.
Mo, et al. Expires 3 April 2027 [Page 9]
Internet-Draft Agent Affinity September 2026
* Decisions that are not reviewable. When the reason for a placement
is not recorded, an operator cannot distinguish a session that was
correctly pinned from one that was pinned by the absence of a
competitive alternative.
4.3. Goals
The goals of introducing state, storage, and compute affinity into
CATS are the following. They are stated as goals rather than as
requirements; the requirements appear in Sections 8, 9, 11, 12, 13,
and 14.
* G1. Make readiness visible. A selection function is able to learn,
for a candidate and for the elements that a step needs, whether the
element is present, at which tier, for how long, and whether the
computation that would consume it is ready.
* G2. Make the alternatives comparable. The cost of reusing what is
present is comparable with the cost of transferring it and with the
cost of reconstructing it, on a common basis and for the same step.
* G3. Keep constraints out of the ranking. Locality, tenancy,
capability, and consistency conditions are evaluated as constraints
before any affinity preference is valued, so that a preference is
never satisfied by paying for it in a dimension in which the step
has a constraint.
* G4. Assure affinity rather than assume it. Affinity is established,
maintained, verified, and released by stated mechanisms, with a
procedure that reserves, commits, and can abort.
* G5. Sustain affinity over a horizon. Affinity is maintained across
the steps of a long session, across pauses and stage transitions,
and is re-evaluated when the assumption behind it changes.
* G6. Share without leaking. Reuse across similar agents and across
the sessions of one tenant is enabled where it is permitted, and is
bounded by scope, isolation, quota, and fairness rules where it is
not.
Mo, et al. Expires 3 April 2027 [Page 10]
Internet-Draft Agent Affinity September 2026
5. Problem Statement
5.1. Storage Is Represented Only as Capacity
The CATS metric definition includes storage among the raw metrics
that may be collected, in the form of available capacity, read
throughput, and write throughput [I-D.ietf-cats-metric-definition].
Those raw metrics describe an instance's storage as a resource that a
workload consumes, in the same way that processor utilization or
bandwidth describe other resources. They do not describe whether a
particular piece of state is available, and they cannot be used to
answer the question that matters for agent service selection.
5.2. Affinity Is Instance-Level
The framework defines service contact instance affinity, which keeps
the traffic of a session on the same instance
[I-D.ietf-cats-framework]. That concept is binary with respect to the
instance: traffic either stays on the instance or does not. For agent
services, affinity has a structure: a session may be able to continue
on a new instance without penalty if one element of its working set
is transferable and cheap, while it must stay where it is if the
element is not transferable or its transfer is prohibited. An
instance-level affinity flag cannot express that distinction.
5.3. Transfer and Re-computation Are Not Comparable
A CATS system that can see storage capacity but not which state is
present cannot compare the two ways of obtaining missing state.
Transferring a KV cache of a given size over a path with a given
capacity and latency has a cost; recomputing it from the request has
a cost that depends on the capability of the instance and the length
of the context. The cheaper option determines which candidate is
preferable, and the current metric set expresses neither.
5.4. Compute Affinity Is Not Represented
Mo, et al. Expires 3 April 2027 [Page 11]
Internet-Draft Agent Affinity September 2026
The computing metrics of the framework describe an instance: its
capability, its utilization, its queue. A step, however, does not
need an instance; it needs a specific computation to be ready.
Nothing in the current metric set states whether the artifact of the
model revision that the step requires is resident at the candidate,
whether the adapter that the session uses is loaded, whether the
precision that the step needs is available on the accelerator of that
candidate, or whether the runtime at the candidate can consume the
state that is available there.
The consequence is that a capability requirement is expressed as a
filter on an instance rather than as a property of the pair of a
requirement and a candidate: a candidate that is capable in general
may still be unable to serve the step without a load, and a load that
takes longer than the step itself is not visible to a selection that
compares only utilization and path quality. The same applies to the
state: a candidate whose runtime cannot consume the KV layout in
which a prefix is stored holds a state that has no reuse value for
that step.
5.5. Reuse across Similar Agents and Tenants Is Not Represented
An agent service usually serves many sessions whose steps require the
same computation and the same reusable material: the same system
prompt prefix, the same tool schemas, the same retrieval corpus, the
same model artifact, the same adapter. Because the current
information is per instance capacity and load, the fact that a prefix
is already resident is invisible to the other sessions that could use
it, and the pre-fill of a shared prefix is paid once per session
instead of once per reuse group. Prefix-level steering for agent
traffic has been proposed on the forwarding side
[I-D.zhang-cats-token-aware-ts]; the corresponding reuse side is not
represented.
The tenant dimension has the same shape. A tenant has a placement
domain, a capacity expectation, and an isolation requirement, and its
sessions share state that may not be shared with another tenant.
Neither the placement domain of a tenant nor the scope of a reusable
element is expressible in the current selection input, and neither is
the effect of affinity on the fairness between tenants.
5.6. The Long Horizon Amplifies Each of These Gaps
Mo, et al. Expires 3 April 2027 [Page 12]
Internet-Draft Agent Affinity September 2026
An agent session can run for a long time, with pauses that exceed the
lifetime of a cache entry, with stages that have different
requirements, and with a working set that grows while it runs. Every
one of the gaps above becomes more expensive over a longer horizon:
state that is lost at a pause costs a resume from an older
checkpoint, a placement that is re-evaluated at every step costs a
transfer per step, and a stage transition that is not anticipated
costs a cold start in the middle of a task. A selection input that is
valid for a single request is therefore not sufficient, and the
assurance procedure for a long session is not the same as the one for
a single step.
6. Working Set Elements Relevant to Agent Services
6.1. Element Classes
The elements below are the ones that agent service selection is
expected to take into account. They differ in size, in volatility, in
whether they may be shared, and in whether they may move.
Element | Volatility | Reuse scope | Locality
-------------------+-------------+---------------+-----------
Session context | grows per | per session | jurisdiction
| turn | |
Retrieved data set | per query | per tenant | policy
| | | dependent
Tool results | short-lived | per session | endpoint
| | | bound
KV cache or prefix | grows per | per session | accelerator
state | token | or per prefix | bound
Plan and scratch | per step | per session | jurisdiction
state | | |
Long-term memory | appended | per tenant or | policy
or index | | per user | dependent
Model artifact | versioned | per site or | hardware
| | per tier | bound
Adapter or | versioned | per model or | license
quantization | | per tenant | bound
artifact | | |
Training | per | per job | jurisdiction
checkpoint | iteration | |
Runtime and kernel | versioned | per site | hardware
image | | | bound
The last four elements arise in long-running jobs, such as the
distributed training use case of [I-D.ietf-cats-usecases-requirements],
and are listed here because a long-running session raises the same
selection question as an agent session.
Mo, et al. Expires 3 April 2027 [Page 13]
Internet-Draft Agent Affinity September 2026
The list is not exhaustive, and it is expected that deployments will
add elements. What matters for this document is that the properties
in the columns are the ones that selection needs, and that they are
not properties of the instance alone but of the pair of the element
and the instance.
The reuse scope column is the basis of Section 14: an element whose
reuse scope is wider than a session is material that several
sessions, several agents, or one tenant can share, and it is the
element class for which the cost of re-establishing the element is
paid most often for no reason.
6.2. Compute-Side Readiness
The compute side has its own working set, and it is as expensive to
re-establish as the state side. It consists of the artifacts and the
runtime properties that a step requires: the weights of the model
revision that the step needs, the adapter or the quantization
artifact that the session uses, the runtime and kernel images that
the accelerator requires, and the compiled or cached forms of the
operations that the step executes.
These elements differ from the state elements in one respect that
matters to the selection: they are shared by many sessions, they are
versioned rather than volatile, and their re-establishment cost is a
load rather than a re-computation. A selection that compares two
candidates on the basis of free accelerator capacity alone will
prefer the candidate that has the memory free and will then pay the
load, which is exactly the cost that a compute affinity view makes
visible.
6.3. Why the Properties Are Pairwise
Each property in the table above is a property of a pair: the element
and the instance. The tuple of the model revision, the tokenizer, and
the precision that produced a KV state is part of the identity of
that state for the purpose of reuse; the runtime of a candidate
determines whether the state that the candidate holds is usable by a
step; and the free capacity of a candidate determines whether the
element fits. A summary of an instance that is not expressed against
the element that a step needs therefore cannot answer the question
that selection asks.
7. State, Storage, and Compute Affinity
7.1. Affinity, Preference, and Locality
Mo, et al. Expires 3 April 2027 [Page 14]
Internet-Draft Agent Affinity September 2026
Affinity is a preference, not a constraint. It states that a
candidate is better because the material that a step needs is already
in place there, and it may be traded against path quality, load,
cost, and budget. A locality constraint, by contrast, is a rule that
makes a candidate ineligible when the element is not where the rule
permits it to be. The two are frequently confused in designs that
express a locality rule as a very large weight in a score, which
produces a system that admits a violation of the rule when the score
happens to favor it.
The three affinities are defined over different objects and have
different lifetimes.
* State affinity is defined over a working set element and a
candidate. It lasts as long as the element is present at that
candidate, and it is invalidated by eviction, expiry, or a change
of the content of the element.
* Storage affinity is defined over a residency tier, a durable
store, and a candidate. It lasts as long as the element is held
at a tier that the step can use, or as long as the store that can
serve the element is reachable at the cost that the decision、
assumed.
* Compute affinity is defined over a step requirement and a
candidate. It lasts as long as the artifacts and the runtime
properties that the requirement names are resident, and it is
invalidated by a change of the model revision, the adapter,
the precision, or the runtime.
7.2. How the Three Affinities Interact
The three dimensions are not additive, and the reason is that reuse
requires a compatible consumer. The relations below are stated as the
pairwise conditions that a selection has to test.
Pairwise conditions at a candidate C for a step of session S:
Mo, et al. Expires 3 April 2027 [Page 15]
Internet-Draft Agent Affinity September 2026
state_reusable | holds the element AND C can consume it:
| same model revision, tokenizer, and runtime
state_usable | element is at a tier that the step can use, or
| can be promoted into one within the step budget
compute_ready | the artifacts, precision, adapter, and
| accelerator of the step are resident at C
compute_fit | the accelerator has capacity for the state and
| for the working set of the step together
affinity_gain | cost_without_reuse - cost_with_reuse, where the
| two costs are measured for the same step
A candidate that satisfies state affinity but not compute affinity
holds a state it cannot use. A candidate that satisfies compute
affinity but not state affinity must either receive the state or
reconstruct it, and the cheaper of those two is the cost that the
decision has to compare against the gain of the move. A candidate
that satisfies storage affinity without either is close to the
material but not ready to use it, and its value is the difference
between the transfer time from the store and the transfer time from a
distant holder.
The interaction also produces an ordering that the rest of this
document follows: constraints are tested first (Sections 8 and 9),
the cost of not holding the state and of not being compute-ready is
derived next (Section 11), and only then is affinity valued, assured,
and maintained (Sections 12, 13, and 14). This is the same ordering,
and the same division between constraints and quantities, as the
selection mapping of [I-D.mo-cats-agent-selection-mapping].
7.3. The Granularity of an Affinity Decision
Affinity is decided per element and per step, not per instance and
not per session. The practical consequences are as follows.
* A decision may hold for one element and not for another, so the
unit of a decision is the element that a step needs.
* A decision may be revisited at a step boundary without moving the
session, because one element may be re-established where it stands
while the rest of the working set stays in place.
* A decision that covers several steps has a longer lifetime than a
decision that covers one, and the longer lifetime is what makes a
retention commitment worth its capacity.
* A decision that applies to a group of sessions (Section 14) has a
Mo, et al. Expires 3 April 2027 [Page 16]
Internet-Draft Agent Affinity September 2026
lifetime that is bounded by the version of the element, not by the
lifetime of any one session.
8. Requirements and Constraints on the Mapping
8.1. The Mapping Relation
The mapping that this section constrains takes a step of an agent
session and produces, for each candidate instance, the constraints
that the step imposes and the quantities by which the candidates
differ. It is the storage-and-compute part of the mapping described
in [I-D.mo-cats-agent-selection-mapping], stated here at the level of
detail that affinity requires.
The demand side of the mapping is derived from the step and from the
session it belongs to. The supply side is observed at the candidate.
The result is a pairwise valuation.
Mapping element | Item | Content
---------------------+------------------+----------------------
Demand of a step | needs_state | keys, tier, locality,
| | age
| needs_storage | capacity, tier, store
| | locality
| needs_compute | revision, precision,
| | adapter, accelerator,
| | rate
Supply at a | holds_state | key, tier, retention
candidate | |
| offers_storage | capacity, tier,
| | locality
| offers_compute | readiness, free
| | capacity, queue
Pairwise result | state_ready | holds the element and
| | can consume it
| state_cost | min(transfer time,
| | recompute time)
| affinity_gain | cost without reuse,
| | minus the cost with
| | reuse
Three properties of this mapping are required for affinity to be
usable. The mapping is per step, because the requirement of a step is
what a candidate is evaluated against. The mapping distinguishes
constraints from quantities, because a constraint cannot be traded.
And the mapping is recomputable from information that the candidate
can report without disclosing the content of the state, because
otherwise affinity would be expressible only inside a single trust
domain.
8.2. Constraint Families
Mo, et al. Expires 3 April 2027 [Page 17]
Internet-Draft Agent Affinity September 2026
The constraints that affinity imposes on the mapping belong to seven
families. Each of them is evaluated before candidates are ranked, and
each of them can make a candidate ineligible.
* Capacity: the element and the working set of the step fit in the
tier that the step requires at the candidate.
* Capability: the model revision, precision, adapter, and
accelerator class that the step requires are available at the
candidate.
* Compatibility: the candidate can consume the state that it holds
or receives, which requires a matching model, tokenizer, runtime,
and layout.
* Locality: the element and the execution remain within the
jurisdiction, tenant domain, or device boundary that the element
carries.
* Consistency: the element is current with respect to the version
that the step assumes, and is not under an eviction or a replacement
that would make it unusable during the step.
* Isolation: reuse does not cross a session, user, or tenant
boundary that policy protects, and the timing of a hit does not
disclose the presence of another session's state.
* Temporal: the element remains held, and the computation remains
ready,
for the duration that the decision assumes, which for a long-horizon
session is a commitment that exceeds the step.
A budget condition is not a constraint of this kind: exceeding a
budget reduces the value of a candidate rather than making it
ineligible, and it is therefore expressed as a quantity, except where
a session has a hard budget and a step that cannot be served within
it must be failed rather than degraded.
8.3. Mapping Requirements
Mo, et al. Expires 3 April 2027 [Page 18]
Internet-Draft Agent Affinity September 2026
The requirements below are stated using the conventions of BCP 14
[RFC2119] [RFC8174]. They are requirements on the information and on
the mapping that a CATS system needs in order to consider affinity,
and are not protocol requirements.
WM1. A CATS system SHOULD be able to express the demand of a step as
a requirement on each of the state, storage, and compute dimensions,
rather than as a single scalar.
WM2. A CATS system MUST distinguish, in that expression, a constraint
that makes a candidate ineligible from a preference that may be
traded against other quantities.
WM3. A CATS system SHOULD be able to express a state requirement as a
set of required elements, each named by a state handle, with an
indication of whether the element is required to be present at a
stated tier or may be made present at a stated cost.
WM4. A CATS system SHOULD be able to express a storage requirement as
a capacity and a tier requirement, together with a locality
constraint on the durable store that holds the element.
WM5. A CATS system SHOULD be able to express a compute requirement as
a capability requirement, comprising the model revision, the
precision, the accelerator class, the runtime, and any adapter,
together with the rate at which the step needs to consume.
WM6. A CATS system MUST evaluate capacity, capability, compatibility,
locality, consistency, isolation, and temporal constraints before it
values any affinity preference, and MUST NOT satisfy a constraint by
paying for it in another dimension.
WM7. A CATS system SHOULD be able to express the mapping at the
granularity of a step, and SHOULD be able to carry the part of a
session's mapping that is unchanged from one step to the next without
re-deriving it.
WM8. A CATS system SHOULD be able to express the residual requirement
of a step that cannot be served in full, so that a partial reuse, a
partial result, or a degraded execution is expressed as a reduction
of the requirement rather than as a silent failure.
WM9. A CATS system SHOULD be able to state, for each element of the
mapping, the unit in which it is expressed, the component that
observes it, and the bound within which it remains valid.
WM10. A CATS system SHOULD be able to compare the mapping across
candidates of different capability tiers without assuming that the
tiers are interchangeable, and SHOULD preserve the reason when a
candidate is excluded by a constraint.
Mo, et al. Expires 3 April 2027 [Page 19]
Internet-Draft Agent Affinity September 2026
9. Affinity Information Requirements
The requirements below are stated using the conventions of BCP 14
[RFC2119] [RFC8174]. They are requirements on the information that a
CATS system needs in order to consider state, storage, and compute
affinity, and are not protocol requirements.
9.1. State Availability
S1. A CATS system SHOULD be able to determine, for a candidate
instance, whether a working set element is available, and under which
state handle it can be addressed.
S2. State availability information SHOULD be expressed in a form that
can be aggregated, so that an instance is not required to advertise
every element that it holds.
S3. State availability information SHOULD carry an indication of its
freshness, so that a selection function can distinguish a recent
observation from a stale one.
S4. State availability SHOULD be expressed at the granularity at
which reuse is meaningful, which for a KV cache means at the level of
a reusable prefix or block rather than at the level of an entire
instance.
9.2. Residency Tiers and Retrieval Cost
S5. A CATS system SHOULD be able to distinguish the tier at which a
working set element resides, at least between accelerator memory,
host memory, node-local storage, and a shared store.
S6. A CATS system SHOULD be able to represent the cost of making an
element usable at a candidate instance, including both the cost of
transferring it and the cost of recomputing or re-retrieving it.
S7. A CATS system SHOULD be able to identify the cheaper of transfer
and re-computation for a given element and candidate, so that the
selection function can prefer the candidate that yields the lower
total cost.
9.3. Constraints, Consistency, and Sharing
S8. A CATS system MUST be able to express state locality constraints
as constraints that are evaluated before candidates are ranked.
S9. A CATS system SHOULD be able to distinguish state that may be
reused across sessions, across tenants, or not at all, and SHOULD NOT
require a candidate to disclose state that may not be shared.
Mo, et al. Expires 3 April 2027 [Page 20]
Internet-Draft Agent Affinity September 2026
S10. A CATS system SHOULD be able to represent the consistency
conditions under which a stored element may be reused, including at
least whether it is current and whether an eviction is pending.
S11. A CATS system SHOULD be able to represent the cost and the
conditions of changing the instance that serves a session, so that
affinity is applied only when the migration is actually cheaper than
the alternative.
S12. A CATS system SHOULD NOT require the network to learn the
content of a working set element in order to select an instance for
it.
9.4. Compute Affinity and Reuse Scope
S13. A CATS system SHOULD be able to determine, for a candidate,
whether the computation that a step requires is ready there without
being re-established: the model revision, the precision, the adapter,
the runtime, and the accelerator class that the step needs.
S14. A CATS system SHOULD be able to distinguish compute readiness,
which is a property of the pair of a step requirement and a
candidate, from computing load, which is a property of the candidate
alone.
S15. A CATS system SHOULD be able to represent the affinity that a
session, an agent, or a tenant has accumulated with an instance in a
form that survives the end of a turn and the interval between turns.
S16. A CATS system SHOULD be able to represent the reuse scope of an
element, that is, the set of sessions, agents, or tenants for which
the element may be reused, without requiring the content of the
element to be disclosed.
10. State Handles and Their Relationship to CATS Identifiers
Selection requires that a candidate can state which elements it
holds, and that a selection function can compare that statement with
what a step needs. This requires an addressing scheme for working set
elements.
The following properties are proposed for such a handle.
* A state handle identifies a working set element, not a location.
It does not replace the CATS Service Identifier or the CATS Service
Contact Instance ID, and it is not a routable address.
* A state handle is scoped. It is meaningful between the parties
that use it, and it carries or implies the scope within which it
may be used, such as a session, a tenant, a reuse group, or a
site.
Mo, et al. Expires 3 April 2027 [Page 21]
Internet-Draft Agent Affinity September 2026
* A state handle is subject to authorization. Knowledge of a handle
is not authorization to read, transfer, or reuse the state that it
names. Authorization is expected to be provided by the identity and
authorization mechanisms of the environment in which the agent
service runs.
* A state handle is revocable. Eviction, expiry, a change of the
reuse scope, or a policy change can invalidate a handle, and the
selection function is expected to tolerate a handle that has become
invalid.
* A state handle is comparable only within a defined equivalence.
Two handles denote the same reusable state only if the deployment
defines the equivalence, for example equality of a content hash over
a defined model, tokenizer, precision, and prefix.
The last property is the one that the compute and sharing dimensions
make sharper. For the reuse of a model-side state, the equivalence
has to cover the model revision, the tokenizer, the precision, and
the runtime layout, because a KV state that was produced by a
different combination is not reusable by the step even when its
content is identical. This is the same condition that a cache applies
when it decides whether a stored representation may be reused for a
request [RFC9111]: the stored material is reusable only if the key
and the validators of the current request match. The parallel is
stated here because it shows that the condition is a property of the
pair of the material and the consumer, and not a property of the
material alone.
The granularity of a state handle is the granularity at which reuse
is meaningful, as required by S4, and it is the key under which the
storage dimension of agent service selection reports availability
[I-D.mo-cats-agent-selection-mapping].
The framework does not define the syntax of a state handle, and this
document does not propose one. A syntax is needed only if a protocol
is later defined to exchange state availability, at which point the
syntax becomes the subject of the document that defines that
exchange.
11. Measuring Affinity: Metrics and Methods
Mo, et al. Expires 3 April 2027 [Page 22]
Internet-Draft Agent Affinity September 2026
11.1. What Has to Be Measured
Affinity is an expectation about the cost of the next step, and it
can be verified only by measuring what the step actually cost and
what it would have cost without reuse. Four questions therefore have
to be answerable from measurement.
* Is the material there? Presence and residency have to be
observable per reuse key, and not only as an aggregate capacity,
because the value of a candidate depends on the specific element
that a step needs.
* Is the computation ready? Readiness has to be observable as a
property of the pair of a requirement and a candidate, which means
that a measurement of utilization or of installed capability is not
a substitute for it.
* What did the reuse save? Affinity gain has to be measured as a
difference between two costs of the same step, with the basis of the
comparison stated, because a gain measured against a different
baseline is not comparable with a gain measured elsewhere.
* What did affinity cost? The transfer that a decision caused, the
capacity that a retention commitment withheld, and the degradation
that the concentration of sessions caused are costs of affinity and
need to be attributed to the decision that incurred them.
11.2. Affinity Metric Catalogue
The metrics below are the ones that the four questions require. They
are stated at the level of the quantity and its observation point,
not as an encoding, and the levels follow the raw and derived
distinction of [I-D.ietf-cats-metric-definition].
Mo, et al. Expires 3 April 2027 [Page 23]
Internet-Draft Agent Affinity September 2026
Affinity metric | Unit | Level | Observed at
-----------------------+-----------+--------+----------------
Element present under | boolean | raw | C-SMA, per key
a reuse key | | |
Residency tier of an | tier | raw | C-SMA, per key
element | | |
Retention remaining | seconds | raw | C-SMA, per key
Compute readiness of a | boolean | raw | C-SMA, per pair
step requirement | | |
Reuse hit ratio | fraction | derived | C-PS, per key
| | | and window
Reuse distance of a | steps, | derived | C-PS, per
session | seconds | | session
Transfer volume caused | bytes | raw | C-NMA, per
by a decision | | | decision
Transfer time caused | seconds | derived | C-NMA, per
by a decision | | | decision
Recompute volume | tokens, | raw | C-SMA, per step
(re-prefill, reload) | bytes | |
State cost of a step | seconds, | derived | C-PS, per pair
at a candidate | cost | |
Affinity gain of a | seconds, | derived | C-PS, per step
step | cost | |
Selections changed for | count, | derived | C-PS, per
lack of state | reason | | window
Affinity concentration | fraction | derived | C-PS, per site
on an instance | | |
Tenant affinity | bytes, | derived | C-SMA, per
footprint | fraction | | tenant
Reuse refused by scope | count, | derived | C-SMA, per
or policy | reason | | window
The metrics are grouped by the question they answer. Presence,
residency, retention, and readiness answer the first two questions.
Reuse hit ratio, reuse distance, transfer volume and time, recompute
volume, and state cost answer the third. Affinity gain answers the
third and the fourth together, because it states the saving rather
than the cost. Concentration, tenant footprint, and refused reuse
answer the fourth question and are also the inputs of the fairness
rules of Section 14.
11.3. Metric Semantics and Reporting Standards
A quantity that two components report under the same name is
comparable only if the properties below are fixed. They are the
reporting standards that this document asks of a metric definition,
and they follow the treatment of freshness and of unknown values in
[I-D.zhu-cats-metric-semantics].
* Unit and basis. Each metric names the unit in which it is
expressed and the population over which it is computed. A hit
ratio without its window and its population is not comparable
with another hit ratio.
Mo, et al. Expires 3 April 2027 [Page 24]
Internet-Draft Agent Affinity September 2026
* Level. Each metric is stated as raw, observed at a component, or
as derived, computed from raw values. Affinity gain, state cost, and
reuse distance are derived; presence, residency, retention,
readiness, and transfer volume are raw.
* Observation point. Each metric names the component that observes
it and the granularity at which it is observed, which for the
affinity metrics is the reuse key, the pair of a requirement and a
candidate, the session, or the site.
* Freshness. Each reported value carries the time at which it was
observed and the bound within which it may be used, and a value
outside that bound is treated as unknown.
* Percentile and tail. Quantities that describe a distribution are
reported at a stated percentile over a stated window, and a quantity
that is a mean is identified as a mean, so that a tail objective is
not evaluated against an average.
* Unknown. A value that is absent, stale, or of unknown provenance
is reported as unknown rather than as a default, and an unknown value
does not satisfy a constraint.
* Aggregation. Affinity metrics are aggregated per class of element,
per key, per site, and per tenant. Aggregation reduces cardinality but
must not merge populations whose objectives differ, and it must not
turn a per-tenant quantity into a value that discloses another
tenant.
11.4. Measurement Methods
The quantities above are obtainable by four methods, and the choice
among them is a deployment decision.
* Observation at the holder: the instance that holds an element
reports presence, tier, retention, and the outcome of the reuse of the
Mo, et al. Expires 3 April 2027 [Page 25]
Internet-Draft Agent Affinity September 2026
element. This is the most accurate method for the state side and the
only one that can see eviction.
* Observation at the decision point: the component that selects
records the cost of the step for which a decision was taken, the
alternative cost that the same step faced at the candidates that
were rejected, and the reason for the choice. This is the only method
that can produce affinity gain, because the gain is a difference between
two costs of the same step.
* Probing: a periodic or on-demand test of the presence of a key,
used where the holder does not report, and used after a pause to
re-verify the state on which a session depends. Probing has to be rate
limited, because a probe is an observable event and because a probe storm
competes with the traffic it measures.
* Accounting of transfers: the transfer volume and time that a
decision caused, separated from background traffic, so that the cost of
affinity maintenance is visible. This is required for the operational
rule that bounds transfer as a fraction of decisions.
Measurement itself is subject to the constraints of the framework: a
measurement that can say which session produced a byte has to be
protected as session information, and multi-tenant measurement has to
be aggregated so that it does not become an information channel
(Section 18). The OAM functions of [I-D.ietf-cats-oam-fw] are the
natural carrier of these quantities.
11.5. Measurement Requirements
MA1. A CATS system SHOULD define each affinity metric with its unit,
its observation point, and its level, so that two components that
report the same metric report a comparable quantity.
MA2. A CATS system SHOULD measure the presence and the residency of a
working set element per reuse key, and SHOULD NOT require an instance
to enumerate its keys to the network.
MA3. A CATS system SHOULD measure the hit ratio of state reuse per
key and per session over a stated window, and SHOULD report it
together with the window and the population over which it was
computed.
Mo, et al. Expires 3 April 2027 [Page 26]
Internet-Draft Agent Affinity September 2026
MA4. A CATS system SHOULD measure the transfer that a decision causes
and the recomputation that a decision causes separately, so that the
two ways of obtaining missing state can be compared.
MA5. A CATS system SHOULD measure affinity gain as the difference
between the cost of a step that reused state or ready compute and the
cost that the same step would have had without that reuse, and SHOULD
state the basis of the comparison.
MA6. A CATS system SHOULD report latency and cost quantities at a
stated percentile over a stated window, and MUST NOT report a
quantity measured over a population as if it were a property of a
single session.
MA7. A CATS system SHOULD carry the observation time and the validity
bound of each measurement, and SHOULD NOT use a measurement outside
that bound to satisfy a constraint.
MA8. A CATS system SHOULD aggregate affinity measurements per class
of element, per site, and per tenant, and MUST NOT expose a
per-session affinity measurement to a party that is not authorized to
observe that session.
MA9. A CATS system SHOULD distinguish, in its measurements, a miss
that is caused by eviction from a miss that is caused by a scope or
policy prohibition, because the two have different remedies.
MA10. A CATS system SHOULD measure the effect of affinity on the
distribution of load, including the concentration of sessions on the
instances that hold popular state and the effect of an
affinity-driven selection on the tail of other sessions.
12. Affinity Assurance: Mechanisms and Procedures
12.1. Mechanism Catalogue
Affinity is assured by a small set of mechanisms. Each mechanism has
an effect and a cost, and the choice among them is what turns the
information of the previous sections into a placement that is stable
and fair.
Mo, et al. Expires 3 April 2027 [Page 27]
Internet-Draft Agent Affinity September 2026
Mechanism | Effect | Cost or risk
-------------------------+---------------------+--------------
Placement pinning | keeps a session | load
| with the holder of | concentration
| its state |
Retention lease | a holder commits to | capacity
| keep an element | withheld
Replication or warm copy | shortens the path | refresh cost
| to the state |
Prefetch before a stage | state is ready | transfer if
| before the first | unused
| step |
Recompute in place | avoids a transfer | accelerator
| entirely | time
Handoff with reserve and | moves a session | two-instance
commit | without a gap | overlap
Eviction protection | keeps a hot or | capacity for
| shared element | others
| resident |
Admission control | bounds transfers in | queued
| flight | sessions
Quota or reservation | bounds per-session | under-use
| or per-tenant use |
Load-aware override | prevents herding on | affinity gain
| one holder | lost
Degradation ladder | keeps a session | worse
| alive under | objective
| pressure |
Release and eviction | reclaims state that | early loss of
policy | has no horizon | reuse
The mechanisms are complementary and are combined rather than chosen
between. A retention lease without a release policy exhausts the
capacity of the holder; a release policy without a lease makes the
state of a long session disappear during a pause; an override without
a measurement of concentration cannot be applied at the right moment.
12.2. Assurance Procedure
The procedure below is the loop that applies the mechanisms. It is
stated as a sequence of steps with the information that each step
consumes and the decision that it produces, and it is applied at each
decision point of a session.
Mo, et al. Expires 3 April 2027 [Page 28]
Internet-Draft Agent Affinity September 2026
Step | Input | Output
-------+--------------------------+---------------------------
A1 | presence, tier, | per-candidate supply view
| retention, readiness, |
| load |
A2 | supply view, constraints | eligible candidate set
A3 | eligible set, size, | state cost per candidate
| rate, locality |
A4 | gain, path, load, | preferred candidate
| budget, concentration |
A5 | preferred candidate, | commitment or reservation
| next steps |
A6 | missing elements | transfer, prefetch, or
| | reconstruction
A7 | committed state and | verified placement and
| compute | record
A8 | observed gain, decay, | next decision point
| triggers |
Step A2 applies the constraints of Section 8.2 and excludes the
candidates that break one of them, before any affinity is valued.
Step A3 derives the cost of making the working set usable, which is
the cheaper of the transfer and the reconstruction of each missing
element. Step A4 values affinity as the reduction of that cost and
trades it against path quality, load, budget, and the concentration
that the choice would create. Steps A5 and A6 act, and step A7
verifies rather than assumes that the material is in place, because a
commitment can expire between the decision and the step. Step A8
closes the loop and is the reason the procedure is not a one-shot
placement: the observed gain and the decay of the affinity are the
inputs of the next decision.
12.3. Triggers for Re-evaluation
The procedure is re-entered when one of the following events is
observed.
* Session admission, at which the first decision is taken.
* A step boundary, at which the requirement of the next step is
known and may differ from that of the previous one.
* A stage transition of a long-horizon session (Section 13).
* A pause that exceeds the freshness bound of the presence
information that the current placement relies on.
Mo, et al. Expires 3 April 2027 [Page 29]
Internet-Draft Agent Affinity September 2026
* An eviction notice, an expiry of a retention commitment, or a
replacement of a model revision, an adapter, or a runtime that the
session depends on.
* A capacity pressure or a rejection of an admission request, which
indicates that the capacity that the decision assumed is no longer
available.
* A measured breach of the objective of the session, or a measured
concentration of sessions on an instance that degrades the objective
of other sessions.
* A change of the locality, tenancy, or reuse-scope policy that
governs an element of the working set.
12.4. Abort, Failure, and Fallback
A move that has started may fail at any of its steps, and the failure
semantics matter more for affinity than they do for a stateless
decision, because a partially moved working set is worse than either
endpoint. The following properties are expected of the procedure.
* The source of a move remains usable until the target has verified
that it holds the material that the session needs, so that a failure
leaves the session where it was rather than in neither location.
* The verification that a step performs before it runs is the point
at which a failed commitment is detected, and the detection leads to the
degradation classes of the step rather than to an unbounded retry.
* A move that cannot be completed within the affinity budget of the
session is abandoned and the session continues where it is, with the
element re-established locally if that is cheaper than the move.
* When no candidate satisfies the constraints, the session degrades
in the order of the fallback ladder of the selection mapping rather than
failing silently, which is consistent with the notion of a fallback
decision [I-D.pang-cats-fallback-decision-framework].
Mo, et al. Expires 3 April 2027 [Page 30]
Internet-Draft Agent Affinity September 2026
12.5. Assurance Requirements
AM1. A CATS system SHOULD apply affinity as a preference that is
valued after the constraints and alongside load, with a stated weight
or order, and SHOULD NOT allow it to override a constraint.
AM2. A CATS system SHOULD be able to hold a retention commitment for
a working set element for a stated period or until a stated event,
and SHOULD be able to release that commitment before its end.
AM3. A CATS system SHOULD be able to establish the state and the
compute readiness that the next steps need before those steps start,
rather than only at the moment at which a step is served.
AM4. A CATS system SHOULD be able to choose reconstruction in place
as an alternative to a transfer, when the transfer would be larger,
slower, or prohibited.
AM5. A CATS system SHOULD verify, before a step is served, that the
state and the compute readiness that the decision assumed are still
present and usable, and SHOULD fall back when they are not.
AM6. A CATS system SHOULD bound the fraction of decisions that may
trigger a transfer, and SHOULD attribute the transfer that a decision
causes to the session that the decision serves.
AM7. A CATS system SHOULD override affinity when the concentration of
sessions on the instances that hold popular state would degrade the
objective of other sessions, and SHOULD make that override visible in
the selection context.
AM8. A CATS system SHOULD release state that is no longer eligible
for reuse, and SHOULD NOT hold state for a session whose horizon has
ended.
AM9. A CATS system SHOULD re-evaluate an affinity decision when an
assumption behind it changes, including an eviction, an expiry, a
change of model revision or adapter, a change of policy, or a
sustained deviation of the observed gain from the expected gain.
AM10. A CATS system SHOULD be able to abandon a move after it has
started, and SHOULD leave both the source and the target in a usable
state when it does so.
AM11. A CATS system SHOULD apply the same assurance procedure to a
change of residency tier within an instance as to a change of
instance, because both change the cost of the next step.
AM12. A CATS system SHOULD retain the decision and its outcome for
the steps that follow, so that the affinity of a session is not
re-derived from scratch at every step.
Mo, et al. Expires 3 April 2027 [Page 31]
Internet-Draft Agent Affinity September 2026
13. Affinity Assurance for Long-Horizon and Multi-Stage Sessions
13.1. Stage Model
A long-horizon agent session does not have one requirement profile;
it has a sequence of them. A session that plans, then retrieves, then
analyzes, then calls a tool, and then reports has stages whose
working sets overlap in part, whose compute requirements differ, and
whose tolerable latency differs. The stage is the unit at which the
assurance procedure of Section 12 can be applied without paying for a
change at every step.
Stage | Working set that grows | Compute profile
---------------+--------------------------+-------------------
Plan | task state, plan | capable model
Retrieve | retrieved set, index | embedding, search
Analyze | context, intermediate | long context
Act | tool results, arguments | low-latency calls
Report | assembled result | capable model
The model is a description, not a specification: a deployment defines
its own stages. What matters for affinity is that the boundary
between two stages is the point at which the requirement profile
becomes known in advance, and therefore the point at which affinity
can be re-established proactively rather than reactively.
13.2. Phase-Differentiated Assurance
The assurance procedure has a different emphasis in each phase of a
long-horizon session.
* Admission: the session is placed on the basis of its first stage,
and the elements of the later stages that are already known are
prefetched where that is cheap.
* Steady state within a stage: the placement is kept stable, the
retention commitments are renewed as they approach their end,
and the affinity gain is measured rather than assumed.
* Pause: the retention commitment is what protects the session, and
its length is chosen by comparing the cost of holding the state with
the cost of re-establishing it at resume.
Mo, et al. Expires 3 April 2027 [Page 32]
Internet-Draft Agent Affinity September 2026
* Resume: the presence of the elements on which the session depends
is re-verified, and an element that is no longer present is treated
as absent, so that the resume decision is taken on observed rather than
on remembered state.
* Stage transition: the affinity of the next stage is established
proactively, and the placement of the session is changed at that
boundary if the next stage is better served elsewhere.
* Completion: the state is released, and the elements whose reuse
scope is wider than the session (Section 14) are kept where the reuse
group benefits from them.
13.3. Pinning, Retention, and Tier Budget
The capacity that a session may keep warm is finite, and the decision
of what to keep is a comparison of the cost of holding against the
cost of re-establishing. Four rules make that decision tractable.
* Keep the elements whose reconstruction is expensive and whose
reuse is certain, which for a long session is the accumulated context
and the model-side prefix state.
* Keep the elements whose reconstruction is impossible, such as the
result of a tool invocation that cannot be repeated safely or a
retrieval that is no longer reachable.
* Do not keep an element whose reconstruction is cheaper than its
retention, which for a small derived artifact is usually the case.
* Place the elements that are kept at the cheapest tier from which
the step can use them, and promote an element to a faster tier only
for the stage that needs it, because promotion competes with the state
of the steps that are running.
13.4. Checkpoint and Resume after a Pause
Mo, et al. Expires 3 April 2027 [Page 33]
Internet-Draft Agent Affinity September 2026
A pause longer than the retention commitment of the holder, a failure
of the holder, or an administrative move makes the state of a session
unavailable. Three properties make the session resumable.
* A checkpoint of the durable part of the working set, taken at a
stage boundary, bounds the work that a resume can lose.
* The part of the working set that is sufficient to continue is
stated, so that a resume does not attempt to re-establish the whole
set before the session can make progress.
* The resume is a decision point like any other: the presence of the
elements is verified, the cost of the resume is compared across
candidates, and the placement that was in force before the pause is
not assumed to be valid.
13.5. Stage Transitions
A stage transition is the natural point at which a placement may
change, and a change elsewhere is what the stability rule of the
selection mapping is intended to prevent. Three properties apply.
* The requirement profile of the next stage is known at the
boundary, so
the constraint test of the next stage can be performed before the
stage starts, and a candidate that will fail a constraint can be
excluded without waiting for the failure.
* The elements that the next stage needs can be prefetched or
reconstituted during the last part of the current stage, in parallel
with work that is still running, so that the transition does not
begin with a serial load.
* The cost of the transition is attributed to the stage that
benefits from it, and a transition that serves several remaining
stages is charged to the session rather than to one step.
13.6. Drift, Decay, and Re-evaluation
Mo, et al. Expires 3 April 2027 [Page 34]
Internet-Draft Agent Affinity September 2026
Affinity decays, and the decay is gradual rather than binary. The
working set grows with the session, so the element that fits at
admission may not fit later. The context that the session uses may
become less similar to the prefix that is cached. The load of the
holder may grow until the affinity gain no longer compensates for it.
A session whose placement was correct at step ten can be wrong at
step thirty without any single event having changed.
Two quantities make the decay observable: the observed affinity gain,
compared with the gain that the placement was expected to deliver,
and the reuse hit ratio of the session over a recent window. A
sustained reduction of either, beyond the tolerance of the session,
is the trigger for re-evaluation. Re-evaluation is not the same as
movement: it may conclude that re-establishing one element at the
current instance is cheaper than moving the session, which for a
large working set is usually the case.
13.7. Budget Pacing across Stages
A long-horizon session has a budget, and affinity maintenance spends
it. The cost of a stage transition, the capacity that a retention
commitment withholds, and the traffic of a prefetch are all
affordable in isolation and not affordable in a loop. The procedure
therefore paces affinity maintenance over the remaining horizon: at
each stage boundary, the budget that the remaining stages need is
reserved before the affinity of the current stage is extended, and
the transfer that a decision causes is charged to the session budget
that the decision serves.
13.8. Failure and Recovery
The failure modes that matter for affinity over a horizon are the
loss of the holder, the loss of the material, and the loss of the
placement's value. Each has a recovery path, and the path that was
taken is part of the decision record.
* Loss of the holder: the session is resumed from a checkpoint, on
an instance chosen by the same procedure, with the elements that
survived reused where they are reachable.
* Loss of the material: the element is reconstructed if that is
cheaper than re-establishing it elsewhere, and the reconstruction
is measured as a recompute event so that the failure is visible in the
accounting.
* Loss of the value: the placement no longer delivers a gain, and
the session is moved at the next stage boundary rather than immediately,
Mo, et al. Expires 3 April 2027 [Page 35]
Internet-Draft Agent Affinity September 2026
unless a constraint makes the current placement ineligible.
13.9. Multi-Agent Stages
A stage may be served by several agents that share a working set: a
planner and a set of workers, or a set of agents that read the same
retrieved material. Affinity for such a stage is a property of the
shared element as well as of the session, and the placement of the
agents of one stage is decided together where the shared element is
large, because duplicating it across instances costs as much as the
transfer that the shared placement avoids.
13.10. Long-Horizon Requirements
LH1. A CATS system SHOULD represent a long-horizon session as a
sequence of stages, and SHOULD be able to state, for each stage, the
elements and the compute capabilities that the stage needs.
LH2. A CATS system SHOULD keep the placement of a session stable
across the steps of a stage, and SHOULD change it at a stage boundary
rather than within a stage, unless a failure or a violated constraint
forces the change.
LH3. A CATS system SHOULD maintain the affinity of a session across a
pause, and SHOULD carry a retention commitment that covers the
expected idle interval where the cost of losing the state exceeds the
cost of holding it.
LH4. A CATS system MUST re-verify the presence of the elements that a
session depends on when the session resumes, and MUST treat an
element that is no longer present as absent.
LH5. A CATS system SHOULD be able to checkpoint the durable part of a
working set, so that a session can be resumed on another instance
after a failure, an eviction, or an administrative move.
LH6. A CATS system SHOULD state the part of a working set that is
sufficient to continue a session, so that a resume reconstructs only
the material that the session needs.
LH7. A CATS system SHOULD be able to establish the affinity of a
stage before the stage starts, including the prefetch of the
artifacts and the elements that the next stage needs.
LH8. A CATS system SHOULD detect affinity decay, that is, a sustained
reduction of the observed gain of the current placement, and SHOULD
re-evaluate the placement when that reduction exceeds the tolerance
of the session.
Mo, et al. Expires 3 April 2027 [Page 36]
Internet-Draft Agent Affinity September 2026
LH9. A CATS system SHOULD pace the cost of affinity maintenance over
the remaining horizon of a session, and SHOULD NOT spend the budget
of the remaining stages to preserve the state of the current one.
LH10. A CATS system SHOULD attribute the cost of a stage transition
to the stage that benefits from it, and SHOULD attribute a transfer
that serves several remaining stages to the session rather than to a
single step.
LH11. A CATS system SHOULD be able to recover the affinity of a
session after the failure of the instance that holds its state, by
falling back to a checkpoint or to a reconstruction path, and SHOULD
record which path was taken.
LH12. A CATS system SHOULD decide the placement of a multi-agent
stage that shares a working set as one decision, and SHOULD account
for the duplication when the agents of a stage are placed on
different instances.
14. Affinity for Similar Agents and for Tenants
14.1. Similarity and Reuse Groups
Two agents are similar for the purpose of this document when their
steps require the same reusable elements, and not when they are the
same software or serve the same user. The reusable elements of
Section 6 are frequently shared: a system prompt prefix is common to
every session of an agent service, a tool schema is common to every
agent that uses the tool, a retrieval corpus is common to a tenant,
and a model artifact is common to every step of every session that
runs on it.
Affinity for a group has a different economics from affinity for a
single session. The cost of establishing an element is paid once and
is amortized over the reuse group, so an element whose
re-establishment is expensive is worth keeping even when the
individual session that caused it has ended. The risk is also
different: the group is what makes a position on one instance
popular, and it is what makes the isolation and fairness rules of the
following subsections necessary.
14.2. Elements That Can Be Shared
Mo, et al. Expires 3 April 2027 [Page 37]
Internet-Draft Agent Affinity September 2026
Reusable element | Reuse group | Precondition
-----------------------+---------------------+--------------
System prompt prefix | agents of one agent | same model
state | service | and tokenizer
Tool schema and | agents of one tool | same template
templates | set | version
Retrieved corpus and | tenant or user | same rights
index | | and version
Embeddings of shared | tenant | same
documents | | embedding
| | model
Model weights and | site or capability | same
adapter | tier | precision and
| | revision
Policy and evaluation | agent service | same policy
state | | version
Tenant memory and plan | tenant only | no
state | | cross-tenant
| | scope
The conditions of reuse in the third column are what make a group a
reuse group: an element may be reused by the members of the group
only when the conditions hold at the candidate that holds it. A
prefix state produced by one model revision is not reusable by a step
that runs another, even when the prompt text is identical, because
the state is not a copy of the prompt.
14.3. Conditions for Safe Reuse
Shared reuse is safe when the following conditions are met, and each
of them is a condition that the holder evaluates rather than one that
the requester asserts.
* Identity of the computation: the model revision, tokenizer,
precision, runtime, and adapter that produced the element match those
that the reusing step requires.
* Identity of the material: the element is the same version of the
same material, established by a defined equivalence (Section 10) rather
than by a name that two producers may use for different content.
* Authorization: the reuse group of the element includes the session
that requests the reuse, and the holder enforces that membership.
Mo, et al. Expires 3 April 2027 [Page 38]
Internet-Draft Agent Affinity September 2026
* Isolation: the reuse does not disclose the content of one session
to another, and does not disclose the presence of one session's
state to another through the timing of a hit.
* Currency: the element has not been superseded by a version that
the reusing step assumes, and is not under a pending eviction or
replacement that would make it unusable during the step.
The fourth condition is the one that is most easily lost in design,
because sharing is implemented for efficiency and its observability
consequences are considered later. A shared hit that is faster than a
miss is an observable signal, and in a multi-tenant system it is a
signal about another tenant's activity if the sharing is not bounded.
14.4. Tenant Affinity Domains and Isolation
A tenant is the unit at which placement policy, capacity expectation,
and isolation are usually expressed. Three properties apply to the
affinity of a tenant.
* A tenant has an affinity domain, that is, the set of instances at
which its sessions and its state may be placed. The domain is a
constraint: a candidate outside the domain is ineligible for that
tenant's sessions and for the state that belongs to the tenant,
regardless of the affinity gain that it would produce.
* The scope of a reusable element is part of the element, not a
property of the requester. An element that a tenant holds is reusable
by the sessions of that tenant within the domain, and it is not reusable
outside it, even when the content would be identical for both
tenants.
* The affinity of one tenant is not visible to another. A tenant may
observe its own affinity, and the operator may observe the aggregate,
but neither may observe the presence or the reuse of another tenant's
state.
14.5. Fairness, Herding, and Quota
Affinity concentrates demand, and concentration has to be bounded by
rules that are stated in advance rather than applied when the
concentration has already degraded the service.
Mo, et al. Expires 3 April 2027 [Page 39]
Internet-Draft Agent Affinity September 2026
* Herding. Popular reusable elements attract sessions, and the
attraction is self-reinforcing, because the sessions that are steered
to the holder make the element more valuable there. A selection that
applies affinity without a bound therefore produces a distribution that
is worse than the one it started from in the tail, even when the mean
improves.
* Duplication. When the concentration of a reuse group on one
instance degrades the objective of the sessions that share the element,
the remedy is to admit a second copy of the element at another instance
and to split the group, at the cost of establishing the element
twice. The decision to duplicate is a comparison of the two costs,
not a default.
* Quota. A tenant with a large state footprint can occupy the
capacity
that other tenants need. The affinity of a tenant is therefore
bounded by a quota on the capacity that its state may occupy and by a
share of the reuse capacity of a shared instance, and the binding of
the quota is observable rather than silent.
* Fairness of measurement. The affinity gain of one tenant is not
allowed to be achieved by degrading the tail of another, which requires
that the objective of a session be evaluated per tenant and per
communication mode rather than over the aggregate.
14.6. Measurement and Observability across Tenants
The affinity metrics of Section 11 are reported per tenant as well as
per key and per site, and the following restrictions apply to them.
* A tenant may see its own footprint, its own hit ratio, and the
binding of its own quota.
* The operator may see the aggregate footprint, the concentration on
an instance, and the refused reuse, because those are the quantities
that the fairness rules bind to.
Mo, et al. Expires 3 April 2027 [Page 40]
Internet-Draft Agent Affinity September 2026
* Neither may see the per-session affinity of another tenant, and
neither may infer it from a difference in latency, so the counters
that are exposed are aggregated over the population of the tenant
rather than over a key that another tenant could probe.
14.7. Requirements for Similar Agents and Tenants
ST1. A CATS system SHOULD be able to identify the group of sessions,
agents, or tenants that may reuse a given working set element, and
SHOULD be able to express that group as part of the state
information.
ST2. A CATS system SHOULD be able to determine whether two agents may
share an element by comparing the conditions of reuse, including the
model revision, the tokenizer, the runtime, the precision, the
template version, and the policy version, rather than by comparing
their content or their identity.
ST3. A CATS system SHOULD treat a shared element as available for a
step only when the conditions of reuse are satisfied at the candidate
that holds it.
ST4. A CATS system SHOULD be able to express the affinity domain of a
tenant, that is, the set of instances at which the sessions and the
state of the tenant may be placed, and MUST enforce that domain as a
constraint.
ST5. A CATS system MUST NOT reuse a working set element across
tenants, or across sessions of different users within a tenant,
unless the policy of the holder of the element permits that reuse.
ST6. A CATS system SHOULD account for the cost of a shared element
once for the reuse group that establishes it, and SHOULD distinguish
a shared reuse from a private one in its measurements.
ST7. A CATS system SHOULD bound the capacity that the state of one
tenant may occupy at a shared instance, and SHOULD make the binding
of that bound observable.
ST8. A CATS system SHOULD detect the concentration of a reuse group
on a small number of instances, and SHOULD be able to admit an
additional copy of an element when that concentration degrades the
objective of the sessions that share the element.
ST9. A CATS system SHOULD be able to state, per tenant, whether
affinity may be traded against load, distance, or cost, and SHOULD
apply the affinity policy of the tenant to the sessions of that
tenant.
Mo, et al. Expires 3 April 2027 [Page 41]
Internet-Draft Agent Affinity September 2026
ST10. A CATS system SHOULD measure the affinity of a tenant without
disclosing the affinity of an individual session to another tenant,
and SHOULD NOT allow a tenant to observe the presence of another
tenant's state.
ST11. A CATS system SHOULD apply the isolation requirements of the
tenant boundary to a shared element as well as to a private one,
including to the timing of a hit and a miss.
ST12. A CATS system SHOULD be able to revoke the reuse scope of an
element when a tenant or a policy changes, without requiring the
content of the element to be disclosed.
15. Interaction with Existing Work
Metrics: The state, storage, and compute information described here
is intended to appear as an additional category or categories in the
metric framework of [I-D.ietf-cats-metric-definition], alongside the
existing computing, communication, and service categories. In the
terms used there, state availability, residency tier, compute
readiness, and retention are naturally raw (Level 0) metrics, while
the comparison of transfer against recomputation, affinity gain, and
the concentration of affinity are derived (Level 1) quantities.
Data model: If affinity information is to be configured or monitored,
the data model [I-D.ietf-cats-data-model] needs to accommodate it.
The minimum set of objects is a state handle or key space, a
residency tier, a readiness descriptor for a computation, a reuse
scope, a retention commitment, and a policy that states the permitted
scope of reuse.
OAM: The operational indicators for affinity-based selection are the
hit ratio of state reuse, the volume and the latency of the state
that a decision transfers, the affinity gain, the concentration of
sessions on the instances that hold popular state, and the number of
selections changed because state was not available where it was
expected. These are candidates for the monitoring functions of
[I-D.ietf-cats-oam-fw].
Selection mapping: The mapping described in
[I-D.mo-cats-agent-selection-mapping] provides the framework in which
the requirements of Sections 8 to 14 are applied: the descriptor of a
step, the resource view of a candidate, the decision points, and the
stability requirement are defined there, and this document adds the
state, storage, and compute affinity content of each.
Mo, et al. Expires 3 April 2027 [Page 42]
Internet-Draft Agent Affinity September 2026
Distributed cache and storage protocols: Moving working set elements
between instances is a data transfer problem that is expected to use
existing transport building blocks. This document does not define a
transfer protocol, and multiple such protocols may be used in one
deployment. The reuse condition of Section 10 is the same condition
that a cache applies when it decides whether a stored representation
may be reused [RFC9111], and cache-control semantics are a useful
model for the retention commitments of Section 12.
15.1. Traceability to the Agent Service Requirements
The requirements of [I-D.mo-cats-agent-service-characteristics] are
addressed as follows. The last two rows map the characteristics of
that document that this document develops further.
* R1, R2: WM1, WM2, WM6.
* R8, R9, R10: S13, S14, WM5.
* R11: S1, S2, S4.
* R12: S6, S7, WM3.
* R13: S8, S9, WM6.
* R14: S15, S16, AM1, ST1, ST3.
* R15, R16: AM6, LH9, LH10.
* R17: MA6.
* R19: AM3, AM4, LH5, LH6.
* R20: AM4, WM8.
* R21: MA4, MA5, MA9.
* R22: MA8, MA10.
* Long-horizon state and memory persistence (Section 5.8 of that
document): LH1 to LH12, AM2, AM8.
* Locality, governance, and tenancy (Section 5.9 of that document):
S8, S9, ST4, ST5, ST11, ST12.
16. Operational Considerations
Mo, et al. Expires 3 April 2027 [Page 43]
Internet-Draft Agent Affinity September 2026
State-based selection changes the shape of the traffic that an
operator sees. State transfer is additional traffic between service
sites, it can be large, and it can be triggered by a steering
decision, which means that a selection that ignores transfer cost can
create the congestion that then degrades the next selection.
Operators are therefore expected to bound the fraction of decisions
that are allowed to trigger a transfer, and to treat the transfer
network as a resource that the selection function must account for.
Affinity can also concentrate load and can increase the impact of a
failure: the instance that holds the state of many sessions is a
higher-value target and a higher-impact failure than an instance that
holds none. Deployments are expected to decide how much state is
worth keeping warm, and to have a fallback path when the preferred
instance is unavailable, which is consistent with the notion of a
fallback decision [I-D.pang-cats-fallback-decision-framework].
Five operational properties follow from the mechanisms of Section 12.
* The retention policy of an instance is a capacity decision, not
only a performance decision, because a commitment that is not released
withholds capacity from the steps that are running.
* The override rule has to be testable in advance. An operator is
expected to know at which measured concentration the override will
be applied, rather than discovering it during an incident.
* The failure of the holder of a large amount of state is a
common-mode event for every session that it holds. A deployment is
expected to bound the amount of state that one instance holds for
one tenant, and to keep the recovery path (checkpoint or reconstruction)
exercised.
* Affinity metrics have to be attributable to a decision, because a
hit ratio that cannot be attributed to a placement cannot be used to
correct the placement.
* The duplication decision of Section 14 has a cost that appears in
the storage metrics and a benefit that appears in the tail of the
objectives. An operator is expected to state which of the two is
bounded, so that the decision can be taken consistently.
Mo, et al. Expires 3 April 2027 [Page 44]
Internet-Draft Agent Affinity September 2026
17. Security Considerations
State is more sensitive than capacity. Exposing which elements an
instance holds reveals patterns of use even when the content is not
exposed, and the ability to request a transfer is the ability to move
data between locations. The following considerations follow.
Content protection: State that is transferred between instances
should be protected in the same way as the session data from which it
is derived, including at rest where the tier is persistent.
Handle confidentiality and unforgeability: A state handle that can be
guessed or forged is a handle that can be used to probe for the
existence of state, and possession of a handle MUST NOT grant access
to the state. Handles should be unguessable, scoped, and validated
against authorization before use. A handle that names a shared
element is scoped to the reuse group of that element, so that
knowledge of it does not disclose the existence of state outside the
group.
Cross-tenant leakage: Reuse of state across sessions or tenants
creates a direct path to information disclosure, including through
the timing of a hit or a miss. Sharing scope should be enforced by
the entity that holds the state, not only by the entity that requests
it.
Poisoning: A participant that can cause incorrect state to be
associated with a valid handle can influence the output of the
sessions that reuse it. Integrity of the association between a handle
and its content should be established by the holder of the state.
Denial through eviction: Because state availability is a resource
that can be exhausted, a participant that can cause eviction can
degrade other sessions. Implementations should not allow one session
to evict the state of another without authorization. Affinity widens
this surface, because an element that is shared by a reuse group is a
single object whose eviction degrades the whole group, and because
the retention commitments of one session compete with the state of
the others.
Compute substitution: A selection that treats capability tiers as
interchangeable can place a step on a candidate that is cheaper but
less capable, which changes the result of the step rather than only
its cost. The constraint families of Section 8.2 keep capability out
of the affinity valuation for this reason.
18. Privacy Considerations
Mo, et al. Expires 3 April 2027 [Page 45]
Internet-Draft Agent Affinity September 2026
Working set elements are derived from user inputs and therefore
inherit their sensitivity. Even aggregated state availability
information can reveal activity, and transfer events reveal which
sessions are active between which sites. Where such information is
exposed to the network it should be minimized, aggregated where
possible, and retained only as long as it serves the selection
function.
Locality constraints are frequently imposed because of legal or
regulatory requirements, and where a constraint is expressed in the
information exchanged, the accuracy of that expression determines
whether the requirement is met. Implementations SHOULD treat an
unknown constraint as a prohibition rather than as permission.
Shared reuse adds two considerations. First, a reuse group is an
inference surface: a party that can observe the hit ratio of a shared
element learns how many sessions use it and when, even when it learns
nothing about their content. Second, the revocation of a reuse scope
is a privacy control, and it is expected to be effective for the
state that is already held and not only for the state that is
subsequently created. The metrics of Section 11 are therefore
reported per tenant and aggregated over a population, and not per key
where a key could be probed by another tenant.
19. IANA Considerations
This document has no IANA actions.
20. Normative References
[I-D.ietf-cats-framework] Li, C., Du, Z., Boucadair, M., Contreras,
L. M., et al., "A Framework for Computing-Aware Traffic Steering
(CATS)", Work in Progress, Internet-Draft,
draft-ietf-cats-framework-24, September 2026.
[I-D.ietf-cats-metric-definition] Yao, K., et al., "CATS Metrics
Definition", Work in Progress, Internet-Draft,
draft-ietf-cats-metric-definition-12, September 2026.
[I-D.mo-cats-agent-service-characteristics] Mo, Y., Yang, D., Zhou,
C., "AI Agent Service Characteristics and Their Implications for
Computing-Aware Traffic Steering", Work in Progress, Internet-Draft,
draft-mo-cats-agent-service-characteristics-00, September 2026.
21. Informative References
[CATS-CHARTER] IETF, "Computing-Aware Traffic Steering (CATS) Working
Group Charter",
.
Mo, et al. Expires 3 April 2027 [Page 46]
Internet-Draft Agent Affinity September 2026
[I-D.ietf-cats-usecases-requirements] Yao, K., et al.,
"Computing-Aware Traffic Steering (CATS) Problem Statement, Use
Cases, and Requirements", Work in Progress, Internet-Draft,
draft-ietf-cats-usecases-requirements-14, September 2026.
[I-D.mo-cats-agent-selection-mapping] Mo, Y., Yang, D., Zhou, C., "A
Selection Mapping Framework for AI Agent Services in Computing-Aware
Traffic Steering", Work in Progress, Internet-Draft,
draft-mo-cats-agent-selection-mapping-00, September 2026.
[I-D.ietf-cats-data-model] Yao, H., Lin, C., et al., "Data Model for
Computing-Aware Traffic Steering (CATS)", Work in Progress,
draft-ietf-cats-data-model-00, September 2026.
[I-D.ietf-cats-oam-fw] Fu, H., Xiong, Q., Du, Z., et al., "Computing-
Aware Traffic Steering (CATS) Operations, Administration, and
Maintenance (OAM) Framework", Work in Progress, Internet-Draft,
draft-ietf-cats-oam-fw-01, July 2026.
[I-D.li-cats-kv-cache-distribution] Li, Z., et al., "KV Cache
Distribution for Distributed LLM Inference: Use Case and
Requirements", Work in Progress,
draft-li-cats-kv-cache-distribution-00, July 2026.
[I-D.pang-cats-fallback-decision-framework] Pang, R., Ed., Han, M.,
Ed., Huang, T., Ed., "CATS Fallback Decision Framework", Work in
Progress, Internet-Draft,
draft-pang-cats-fallback-decision-framework-00, July 2026.
[I-D.zhang-cats-token-aware-ts] Zhang, N., Ed., Han, M., Ed., Yi, X.,
Ed., "A token-aware traffic steering solution for agent service",
Work in Progress, draft-zhang-cats-token-aware-ts-00, March 2026.
[I-D.zhu-cats-metric-semantics] Zhu, M., "Operational Semantics for
CATS Metric Consumption", Work in Progress,
draft-zhu-cats-metric-semantics-01, August 2026.
[RFC2119] Bradner, S., "Key words for use in RFCs to Indicate
Requirement Levels", BCP 14, RFC 2119, DOI 10.17487/RFC2119, March
1997, .
[RFC8174] Leiba, B., "Ambiguity of Uppercase vs Lowercase in RFC 2119
Key Words", BCP 14, RFC 8174, DOI 10.17487/RFC8174, May 2017,
.
[RFC9111] Fielding, R., Nottingham, M., Reschke, J., "HTTP Caching",
STD 98, RFC 9111, DOI 10.17487/RFC9111, June 2022,
.
Acknowledgments
Mo, et al. Expires 3 April 2027 [Page 47]
Internet-Draft Agent Affinity September 2026
The authors would like to thank the participants of the CATS working
group for the discussions that shaped this document.
Authors' Addresses
Y. Mo
Huazhong University of Science and Technology
Email: moyj@hust.edu.cn
D. Yang
Huazhong University of Science and Technology
Email: d202581903@hust.edu.cn
C. Zhou
Huazhong University of Science and Technology
Email: m202474228@hust.edu.cn
Mo, et al. Expires 3 April 2027 [Page 48]