CATS Y. Mo Internet-Draft Huazhong University of Science and Technology Intended status: Informational 30 September 2026 Expires: 23 April 2027 AI Agent Service Characteristics and Their Implications for Computing-Aware Traffic Steering draft-mo-cats-agent-service-characteristics-00 Abstract AI agent services place a new class of demands on the network: sessions are long-lived and stateful, a single user request expands into multiple model invocations and tool calls, and the quality of the result depends jointly on the forwarding path, on the computing capability that is available at the selected site, and on whether the state that the agent needs is already present near that site. Computing-Aware Traffic Steering (CATS) already selects a service contact instance using a combination of computing and network metrics, but the CATS framework, metric, and data model documents were not written with agent services in mind. The document covers long-horizon tasks and their long-tail behaviour, discrete and distributed tool invocation, the cost of context growth and compression, the communication modes used by agent services, and the persistence of session state and memory. It reports measured distributions of model, state, and software artifact sizes, states what a CATS system must measure in order to steer this traffic, and identifies the information that a CATS system needs in order to select an instance for an agent service. This document does not define any protocol extension, metric encoding, or data model. Status of This Memo This Internet-Draft is submitted in full conformance with the provisions of BCP 78 and BCP 79. Internet-Drafts are working documents of the Internet Engineering Task Force (IETF). Note that other groups may also distribute working documents as Internet-Drafts. The list of current Internet- Drafts is at https://datatracker.ietf.org/drafts/current/. Internet-Drafts are draft documents valid for a maximum of six months and may be updated, replaced, or obsoleted by other documents at any time. It is inappropriate to use Internet-Drafts as reference material or to cite them other than as "work in progress." This Internet-Draft will expire on 23 April 2027. Copyright Notice Y. Mo Expires 23 April 2027 [Page 1] Internet-Draft Agent Service Characteristics October 2026 Copyright (c) 2026 IETF Trust and the persons identified as the document authors. All rights reserved. This document is subject to BCP 78 and the IETF Trust's Legal Provisions Relating to IETF Documents (https://trustee.ietf.org/ license-info) in effect on the date of publication of this document. Please review these documents carefully, as they describe your rights and restrictions with respect to this document. Code Components extracted from this document must include Revised BSD License text as described in Section 4.e of the Trust Legal Provisions and are provided without warranty as described in the Revised BSD License. Discussion Venues Discussion of this document takes place on the CATS Working Group mailing list (cats@ietf.org), which is archived at https://mailarchive.ietf.org/arch/browse/cats/. Table of Contents Y. Mo Expires 23 April 2027 [Page 2] Internet-Draft Agent Service Characteristics October 2026 1. Introduction . . . . . . . . . . . . . . . . . . . . . . . . . 3 2. Conventions and Definitions . . . . . . . . . . . . . . . . . .4 3. Terminology . . . . . . . . . . . . . . . . . . . . . . . . . .5 4. Problem Statement . . . . . . . . . . . . . . . . . . . . . . .5 5. Agent Service Characteristics . . . . . . . . . . . . . . . . .6 5.1. Session-Oriented and Stateful . . . . . . . . . . . . . . 6 5.2. Multi-Stage Task Execution and Discrete Tool Invocation . 7 5.3. Long-Horizon Tasks and Long-Tail Behaviour . . . . . . . .8 5.4. Context Growth, Compression, and Their Resource Costs . . 9 5.5. Heterogeneous and Tiered Model Capability . . . . . . . .10 5.6. Agent Communication Modes and Their Network Requirements . . . . . . . . . . . . . . . . . . . . . . . .10 5.7. Long-Horizon State and Memory Persistence . . . . . . . .12 5.8. Locality, Governance, and Tenancy Constraints on State . 13 5.9. Cost, Energy, and Token Budget Sensitivity . . . . . . . 13 5.10. Turn-Level and Task-Level Performance Objectives . . . .14 6. Observed Distributions of Transfer, Storage, and Compute Resources . . . . . . . . . . . . . . . . . . . . . . . . . . .14 6.1. Model and State Artifacts . . . . . . . . . . . . . . . .14 6.2. Software Artifacts and Protocol Implementations . . . . .16 6.3. Implications for CATS . . . . . . . . . . . . . . . . . .16 7. Measurement Requirements . . . . . . . . . . . . . . . . . . .17 8. Mapping of Agent Service Characteristics to CATS Dimensions . 18 8.1. Forwarding Dimension . . . . . . . . . . . . . . . . . . 18 8.2. Computing Dimension . . . . . . . . . . . . . . . . . . .19 8.3. Storage and State Dimension . . . . . . . . . . . . . . .19 8.4. Characteristic-to-Dimension Matrix . . . . . . . . . . . 20 9. Requirements . . . . . . . . . . . . . . . . . . . . . . . . .20 9.1. General Requirements . . . . . . . . . . . . . . . . . . 20 9.2. Forwarding Requirements . . . . . . . . . . . . . . . . .21 9.3. Computing Requirements . . . . . . . . . . . . . . . . . 21 9.4. Storage and State Requirements . . . . . . . . . . . . . 21 10. Relationship to Existing and Ongoing Work . . . . . . . . . .22 11. Operational Considerations . . . . . . . . . . . . . . . . . 24 12. Security Considerations . . . . . . . . . . . . . . . . . . .24 13. Privacy Considerations . . . . . . . . . . . . . . . . . . . 25 14. IANA Considerations . . . . . . . . . . . . . . . . . . . . .25 15. Normative References . . . . . . . . . . . . . . . . . . . . 25 16. Informative References . . . . . . . . . . . . . . . . . . . 25 Acknowledgments . . . . . . . . . . . . . . . . . . . . . . . . . 27 Authors' Addresses . . . . . . . . . . . . . . . . . . . . . . . .27 1. Introduction Y. Mo Expires 23 April 2027 [Page 3] Internet-Draft Agent Service Characteristics October 2026 Computing-Aware Traffic Steering (CATS) enables a network edge node to select a service contact instance on the basis of both computing and network metrics, and to steer the traffic of a service request towards the selected instance [I-D.ietf-cats-framework]. The problem statement, use cases, and requirements for that selection are documented in [I-D.ietf-cats-usecases-requirements], which already includes a use case for distributed AI training and inference and notes that important inference resources include processor cores and the memory used to store key-values and cached tokens. That use case treats an AI inference request as a single request that is served by one service instance. AI agent services, however, are not single requests. An agent service typically maintains a session, performs multi-step execution on behalf of the user, invokes tools and other agents, and depends on state that accumulates across the steps of a session. Consequently, the suitability of a candidate service contact instance depends on more than the current computing load and the current network metrics of that instance: it also depends on the execution state that the instance already holds, on the cost of moving that state to another instance, and on the cost of recomputing it. This document captures those considerations as agent service characteristics, expresses them in terms of three dimensions that a CATS selection function must consider -- forwarding, computing, and storage -- and derives requirements on a CATS system that is expected to steer traffic for agent services. Section 5 describes the characteristics, Section 6 reports measured distributions of the resources involved, Section 7 states the measurement requirements, Section 8 maps the characteristics onto the forwarding, computing, and storage dimensions, and Section 9 states requirements. The intent is to provide input to the CATS metrics work [I-D.ietf-cats-metric-definition], to the CATS data model [I-D.ietf-cats-data-model], and to future work on selection procedures, without defining new wire formats. This document assumes the deployment model of the CATS framework: a single administrative domain in which a network edge node selects a service contact instance and steers traffic towards it. Multi-domain deployment, any change to the underlay network, and any interface that exposes network or computing conditions to applications are out of scope, and "CATS system" is used below to mean a system with these properties. 2. Conventions and Definitions The key words "MUST", "MUST NOT", "REQUIRED", "SHALL", "SHALL NOT", "SHOULD", "SHOULD NOT", "RECOMMENDED", "NOT RECOMMENDED", "MAY", and "OPTIONAL" in this document are to be interpreted as described in BCP 14 [RFC2119] [RFC8174] when, and only when, they appear in all capitals, as shown here. Y. Mo Expires 23 April 2027 [Page 4] Internet-Draft Agent Service Characteristics October 2026 In this document these key words are used to state requirements on the design of a CATS system that supports agent services. They do not describe protocol behaviour, and this document defines no protocol, message format, or data model. 3. Terminology This document uses the terms defined in [I-D.ietf-cats-framework] and [I-D.ietf-cats-usecases-requirements]. The following additional terms are used. Agent: A software entity that performs tasks on behalf of a user or another agent, typically by invoking one or more models, tools, and other agents. Agent Service: A service, identified by a CATS Service Identifier (CS-ID), that provides the execution environment for one or more agents. Agent Session: The long-lived association between a client and an agent service, spanning multiple agent turns. A session is associated with state that is accumulated across turns. Agent Turn: One request-response cycle within an agent session, which may itself consist of several model invocations and tool invocations. Step: One unit of work within a turn, such as a model invocation, a tool invocation, or a retrieval operation. Agent Gateway: A functional element that terminates the agent session towards the client and forwards steps towards serving instances. An agent gateway may act as, or be co-located with, a CATS Traffic Classifier. Working Set: The collection of state that a session needs in order to continue without recomputation or re-retrieval. The working set may include conversation history, retrieved documents, tool results, intermediate model state such as a key-value (KV) cache, and loaded model artifacts. Model Capability Tier: A category that describes the capability of a model that can serve a step, for example a small model co-located with the client, a mid-size regional model, or a frontier model in a central facility. 4. Problem Statement A CATS system that steers agent traffic using the metrics and identifiers defined today faces three problems. Y. Mo Expires 23 April 2027 [Page 5] Internet-Draft Agent Service Characteristics October 2026 First, the unit of selection does not match the unit of demand. The CATS identifiers are the CS-ID, which identifies a service, and the CATS Service Contact Instance ID (CSCI-ID), which identifies a service contact instance [I-D.ietf-cats-framework]. An agent service, however, is demanded at the granularity of a session, a turn, or a step, and different steps of the same session may have different requirements: a planning step may need a capable model and tolerant latency, while a tool invocation may need a specific data locality and tight latency. A single selection made at the start of a session does not remain suitable for the whole session. Second, the storage dimension is not represented. Agent services are strongly stateful. Conversation history, retrieved context, tool results, and KV caches determine whether a step can be served quickly at an instance. Existing CATS metrics expose computing capacity and utilization and communication performance, but they do not expose whether reusable state is present at a candidate instance, nor the cost of obtaining it. This gap has been identified for the specific case of a KV cache in [I-D.li-cats-kv-cache-distribution]; the same gap applies to every other component of an agent working set. Third, the relationship between the client's request and the required resources is not expressible. In the current model, a service request is described by the CS-ID and, where an individual proposal is applied, by constraints carried in request packets [I-D.zhang-cats-clients-request-packet]. For agent services, the steering decision would benefit from information that the client or the gateway already knows, such as the number of steps that remain in a planned task, the size of the context that must be carried, the tools that will be invoked, the model capability tier that the task requires, and the phase of the task. Several individual documents have proposed to expose parts of this information, including token features [I-D.zhang-cats-token-aware-ts], tensor and flow semantics [I-D.li-cats-aisemantic-contract], and task segmentation [I-D.li-cats-task-segmentation-framework]. No consolidated description of agent service characteristics exists, and the CATS charter explicitly places exposure of network and compute conditions to applications out of scope, so the discussion of what the network may receive from the application side must be framed carefully. 5. Agent Service Characteristics This section describes the characteristics of agent services that influence steering. Each characteristic is stated in terms that can be observed or inferred by a CATS component or by an agent gateway acting on behalf of a CATS Traffic Classifier. 5.1. Session-Oriented and Stateful Y. Mo Expires 23 April 2027 [Page 6] Internet-Draft Agent Service Characteristics October 2026 An agent session persists across many client interactions and accumulates state. The state includes the conversation history, the working set that has been retrieved for the session, the results of tool invocations, and the model-side state that allows previous context to be reused. The consequence for steering is that the best instance for the next turn is frequently the instance that already holds the session state, even if another instance has more free computing capacity or a better-connected path. Conversely, if the session state can be reproduced cheaply, an instance without the state may be preferred because it is closer or less loaded. State therefore behaves as a first-class selection input, and the cost of not having it locally -- either the cost of transferring it or the cost of recomputing it -- must be comparable with the costs that are already considered. 5.2. Multi-Stage Task Execution and Discrete Tool Invocation A single user request is expanded by the agent into a sequence of steps. Typical steps are planning, retrieval, tool invocation, model invocation, validation, and answer generation. The sequence is data-dependent: the number and identity of the steps are not known when the session starts. Tool invocation deserves separate treatment because it changes the shape of the traffic and of the demand. A single agent step may issue more than one tool call, and in a long-running session the number of tool calls is comparable to or greater than the number of model invocations. Tool calls are therefore not an occasional side effect of agent execution; they are a principal source of traffic and of latency. The following properties of tool invocation matter for steering. * Tool calls are discrete and individually scheduled. Each call is a separate interaction with a separate endpoint, and the endpoints are frequently outside the administrative domain of the service instance. Latency is the sum of the model's decision to call, the round trip to the endpoint, the execution time of the tool, and the return of the result into the context. * Tool calls are bursty and fan out. One turn can issue several calls in parallel, and each may target a different endpoint, so the set of destination prefixes for a single turn is not known in advance. Y. Mo Expires 23 April 2027 [Page 7] Internet-Draft Agent Service Characteristics October 2026 * Tool call patterns are structured and repetitive. Tool calls are described by schemas that the agent knows before the call is made, and the same tools are invoked repeatedly across turns and across sessions. Repetition makes caching, prefetching, and admission of tool endpoints plausible, but only if the network can distinguish the phases of a session. * Tool latency is heterogeneous and heavy-tailed. Retrieval from a remote index, a code execution sandbox, and a third-party API have different and unrelated latency distributions, and the slowest call in a turn bounds the turn. 5.3. Long-Horizon Tasks and Long-Tail Behaviour Agent tasks are long-horizon: they run for many steps, they may pause for minutes while a tool executes or a human approves an action, and their resource consumption is heavy-tailed across tasks. Three observations from published measurements shape the network requirement. * Sessions contain long autonomous loops. A single session can run for tens or hundreds of steps without human intervention, and each step can involve several model and tool invocations. Execution is heavyweight and stateful: the non-model components of an agent stack -- retrieval, tool execution, and state management -- are frequently the dominant contributor to latency, not the model invocation itself. What the network must support is therefore not a sequence of independent requests but a long-lived, stateful execution whose bottleneck moves between the model, the tools, and the network. * Sessions are idle for long periods between bursts. An agent routinely sends a request, then waits while a tool runs or a human approves an action, and the pause can last minutes. Where the serving instance keeps reusable state only for a bounded time, a pause of that length is enough for the state to be evicted, so the next request pays the full cost of re-establishing it. Idle duration is therefore a first-class characteristic: it determines whether state is still where it was, and hence whether a steering decision taken at the start of a pause is still valid at its end. * A small fraction of tasks consumes a large fraction of resources. Long sessions with large contexts and short outputs, and sessions that Y. Mo Expires 23 April 2027 [Page 8] Internet-Draft Agent Service Characteristics October 2026 invoke many different tools, coexist with sessions that do neither, and the distribution across sessions is heavily skewed. Objectives that are evaluated on the mean therefore misrepresent the traffic that operators must plan for, and a CATS system that optimizes the average will under-serve exactly the sessions that dominate cost and user-visible latency. 5.4. Context Growth, Compression, and Their Resource Costs The context of an agent session grows with every turn and with every tool result, and the cost of that growth is paid in four places at once: request and response traffic, prefill computation, accelerator memory, and end-to-end latency. Accelerator memory is the sharpest of the four. The key-value (KV) state that a session requires grows linearly with the number of tokens that the session keeps in context. Measured from published model configurations, a 7B-parameter model with grouped-query attention requires on the order of 55 MiB of KV state per 1,000 tokens in half precision; a 72B model requires about 313 MiB per 1,000 tokens. Table 3 in Section 6 gives the values used here. Two consequences follow: at a context of 128,000 tokens, the KV state of a 7B model (about 7.5 GB) is comparable to half of its weights, and the KV state of larger models exceeds the memory of many accelerators that can host them. Frequent access to the KV state dominates data movement during inference, which is why the placement of that state across memory tiers, and across nodes, is itself a scheduling problem. Because context cannot grow without bound, sessions compress. Four families of techniques are used in practice: truncation of old turns, retrieval of selected passages instead of full history, summarization, and token-level or state-level compression, including quantized KV state. Compression must preserve the commitments of a session -- goals, constraints, decisions, tool results, and safety boundaries -- rather than merely reduce length. For steering this matters twice over: compression changes the traffic and the compute that the session will need, and it changes the state that must be kept, moved, or recomputed. Prefix reuse ties the two together. Serving systems cache the processed prefix of a prompt so that a follow-up request that shares it skips most of the prefill computation and is charged at a fraction of the cost of a full request. Keeping such a prefix warm across a pause costs traffic of its own, whether the prefix is replayed to the serving instance or transferred to another one. From the network's point of view, the same effect appears as an alternative between two traffic patterns: re-sending a prefix and reusing remote state, or recomputing the prefix locally. Y. Mo Expires 23 April 2027 [Page 9] Internet-Draft Agent Service Characteristics October 2026 5.5. Heterogeneous and Tiered Model Capability Agent services are frequently composed of models of different capability, cost, and size, and the same service may offer several variants of the same capability. A step that requires reasoning may need a capable model, while classification, routing, or formatting steps may be served by a small model near the client. Forwarding an agent step to an instance that cannot serve the required capability tier produces a wrong result or a fallback, not merely a slower response. It follows that capability is a constraint for selection, not a quantity that can be traded away against latency or load. Computing metrics that are expressed as a single normalized score [I-D.ietf-cats-metric-definition] may be adequate for ranking instances that are all capable of serving the step, but they cannot express the capability constraint itself. The size spread is large: the model artifacts measured for this document range from below 1 GB to about 689 GB (Section 6). A service that can be served by a 0.5B model and a service that requires a 671B model impose requirements on storage, on load time, and on the network that differ by three orders of magnitude. 5.6. Agent Communication Modes and Their Network Requirements Agent interactions use several distinct communication modes, and the mode determines the traffic pattern that the network must carry. This document does not define, profile, or modify any of these protocols; it records the patterns that a CATS system needs to accommodate. The following modes are in use or in standardization. * Tool and resource access. The Model Context Protocol (MCP) exchanges JSON-RPC 2.0 messages between hosts, clients, and servers over either a standard input/output transport or an HTTP-based transport that uses POST and GET requests with optional server-sent events, and it defines resumability and redelivery for a stream that breaks; connections are stateful and start with capability negotiation [MCP]. Its server features include resources, prompts, and tools, and its client features include sampling and elicitation, so a single session may carry both request-response and server-initiated messages [MCP]. * Agent-to-agent delegation. The Agent2Agent (A2A) protocol binds Y. Mo Expires 23 April 2027 [Page 10] Internet-Draft Agent Service Characteristics October 2026 JSON-RPC 2.0 over HTTP(S) and uses server-sent events for streaming, with an agent card for discovery and a task lifecycle that includes subscription to a task and push-notification callbacks to a URL supplied by the client [A2A]. A task is associated with a context identifier that groups the tasks of one conversation, and the specification defines pagination and state filters for task listing [A2A]. Traffic produced by this mode is therefore bidirectional and long-lived: a streaming response, plus asynchronous callbacks to endpoints that the client nominates. * Broader interoperability. Beyond MCP and A2A, other agent interoperability protocols are being defined, including the Agent Communication Protocol (ACP) and the Agent Network Protocol (ANP). The multipart, MIME-typed nature of some of these bindings means that the payload of a single message may mix text, structured data, and binary content, including media, which changes its size distribution compared with text-only request-response traffic. * Interaction with people (A2P). Where an agent interacts with a person, the requirement is interactive latency rather than throughput: for voice-controlled applications, a sub-second response is the threshold at which the interaction feels seamless, while an agentic workflow with tool calling can add seconds of latency. In this mode the network requirement is dominated by the responsiveness of the first and last hop and by the tail of the latency distribution, not by bandwidth. * Capability-directed routing between agents. Extensions to agent protocols propose that routing decisions be taken on the basis of the declared capabilities of the peers, including the media types a peer can consume; the benefit of such routing depends on the receiving agent being able to use what it receives. This is the same shape of problem as instance selection in CATS: a declared capability set matched against a requirement, with an outcome that depends on whether the receiver can use what it receives. The network-relevant differences between the modes are summarized below. Y. Mo Expires 23 April 2027 [Page 11] Internet-Draft Agent Service Characteristics October 2026 Mode | Transport shape | Message mix | Main need -----+-------------------+----------------+-------------- MCP | stateful session; | requests, | session | stdio or HTTP | responses, | continuity | | notifications | A2A | HTTP request- | task messages; | long-lived | response; stream; | stream chunks | bidirectional | push callback | | ACP | RESTful HTTP | multipart, | payload size | | MIME-typed | and type | | (text, media) | A2P | interactive, | small control | tail latency | latency-critical | plus media | and locality | | streams | -----+-------------------+----------------+-------------- 5.7. Long-Horizon State and Memory Persistence Long-horizon tasks require the session state to survive across pauses, across failures, and across changes of serving instance. Four classes of state must be distinguished because their transport requirements differ. * Session state: the conversation history and the agent's working variables. It is small compared with model state, must remain consistent, and is usually tied to a session identifier. * Model-side reusable state: the KV state of a prefix or of one or more agents. It is large and grows with context (Section 6), and it is only useful if the receiving instance has the same model, the same tokenizer, and a compatible runtime. Such state can be persisted in compressed, for example 4-bit, form and reloaded directly into the attention layer, precisely because the alternative -- re-running the prefill -- is expensive: for a session with a context of several thousand tokens, re-establishing the state can cost seconds of accelerator time per attempt. * Retrieved and derived state: documents, embeddings, and index structures that the session has accumulated. These are often shared across sessions and are governed by data-locality rules. * Durable task state: the record of what the task has done, which is needed to resume after a failure or an approval pause. Y. Mo Expires 23 April 2027 [Page 12] Internet-Draft Agent Service Characteristics October 2026 That state is frequently not co-located with the compute that needs it. Prefix and session states may be held on geographically separated accelerator nodes, so a nearby node offers a short path but little reuse, while a more distant node may hold a longer matching prefix at the cost of additional communication and queueing delay. Persistent prefixes also compete for the same accelerator memory as the KV state of active requests. This is the selection problem of CATS expressed in state terms, and it is the reason why transfer capability, transfer cost, and state location belong in the selection input. Transporting this state imposes requirements that are not those of request-response traffic: transfers are bulk and bursty, they are triggered by caching, migration, or repair decisions, they must be resumable because the sender or receiver may be replaced mid-transfer, and their size grows with the session. Section 7 states what needs to be measured about them. 5.8. Locality, Governance, and Tenancy Constraints on State Agent working sets frequently contain data that carries processing constraints: personal conversation history, enterprise-confidential retrieved documents, or regulated personal data. Two constraints follow. First, the state may not be allowed to move. Steering the next turn of a session to a distant instance may require transferring state that policy prohibits from leaving a jurisdiction, a tenant boundary, or a device. The selection function must be able to treat state location as a constraint rather than as a cost. Second, the state must be isolated between sessions and tenants. An instance that holds reusable state for other tenants is not necessarily a valid target, even if it is technically capable of holding this session's state. Measurement of state reuse must respect the same boundary, because reuse itself leaks information: a cache hit is faster than a miss, and in a shared system that difference is observable, so an attacker may be able to infer information about another user's content from hit and miss patterns 5.9. Cost, Energy, and Token Budget Sensitivity Agent sessions have a monetary and energy cost that is proportional to the work performed, and users or operators frequently impose budgets on that cost. A steering decision that minimizes latency for every step can violate the budget of the session, while a decision that minimizes cost can violate the latency objective of interactive steps. For the network, the relevant part of this characteristic is that cost and energy need to be comparable with latency and load, and that the comparison is per session or per turn rather than per packet. Y. Mo Expires 23 April 2027 [Page 13] Internet-Draft Agent Service Characteristics October 2026 5.10. Turn-Level and Task-Level Performance Objectives The performance objective of an agent service is expressed over the task and its turns rather than over a single packet stream. Objectives that appear in the relevant documents include the latency of the first token and of subsequent tokens for a model invocation, the completion time of a tool invocation, and the completion time of the whole task [I-D.ietf-cats-usecases-requirements] [I-D.li-cats-kv-cache-distribution]. The objective depends on the communication mode. A person-facing session imposes a sub-second budget on the interactive part, a streaming agent-to-agent response is judged on its time to first chunk and its inter-chunk gaps, and a background tool call may be judged on completion time alone. A single latency objective applied to all traffic of a session therefore misallocates resources; the objective needs to be attached to the mode and to the step. Because the objective is defined at task level while steering acts on streams, a CATS system needs to know which stream belongs to which turn and which turn belongs to which task, at least well enough to attribute budget and to avoid optimizing a step at the expense of the task. 6. Observed Distributions of Transfer, Storage, and Compute Resources Selection decisions are comparisons, so the sizes and costs that a selection function compares need to be known at least in order of magnitude. This section records measured distributions of the artifacts and of the state that agent services move, keep, and compute. The numbers are indicative rather than normative: they were measured at a point in time, they will change, and they are given here to establish the relative scale of the three dimensions. Two sources are used. Model and state sizes were measured from the public API of a model registry mirror on 26 September 2026 [MODELREG] and computed from the published configurations of the models concerned; software artifact sizes were measured from a public code hosting API on the same date [GHAPI]. Behavioural properties of agent workloads are described in Section 5 and are stated qualitatively. Each table states the method, so that the values can be reproduced. 6.1. Model and State Artifacts Table 1 gives the size of the weight artifacts of a set of models that are plausibly offered as agent capabilities, from an edge tier to a frontier model. Y. Mo Expires 23 April 2027 [Page 14] Internet-Draft Agent Service Characteristics October 2026 Artifact | GB | Class ---------------------------+--------+-------------- all-MiniLM-L6-v2 | 0.88 | embedding Qwen2.5-0.5B-Instruct | 0.99 | edge bge-m3 | 2.27 | embedding Llama-3.2-1B-Instruct | 2.47 | edge, 4-bit Qwen2.5-7B-Instruct-AWQ | 5.57 | mid Qwen2.5-7B-Instruct | 15.23 | mid Meta-Llama-3-8B-Instruct | 16.06 | multi-format whisper-large-v3 | 24.70 | multi-quant Qwen2.5-7B-Instruct-GGUF | 56.29 | large Qwen2.5-32B-Instruct | 65.53 | MoE Mixtral-8x7B-Instruct-v0.1 | 93.41 | large Llama-3.1-70B-Instruct | 141.11 | large Qwen2.5-72B-Instruct | 145.41 | frontier, MoE ---------------------------+--------+-------------- n = 14; min 0.88 GB; median 20.38 GB; max 688.59 GB Two properties of this distribution matter for steering. First, it spans three orders of magnitude, so a rule such as "the instance with the most free capacity" is not meaningful across services. Second, the artifact that must be loaded before an instance can serve is not the artifact that the client sends: the transfer triggered by a steering decision is a bulk transfer whose duration is measured in seconds to minutes, not in milliseconds. Quantization moves the artifact along that distribution: for the same 7B model, the 4-bit variant measures 5.57 GB against 15.23 GB in half precision. Quantization therefore changes which instances can host a model, how long it takes to load, and which precision trade-offs are acceptable for a step -- all of which a selection function may have to compare. Table 2 gives the size of the KV state that a session accumulates, which is the part of the state that grows without the model changing. Model | MiB/1k tok | KV@8K | KV@32K | KV@128K -----------------+------------+-------+--------+-------- Qwen2.5-0.5B-Ins | 11.72 | 0.10 | 0.40 | 1.61 Qwen2.5-7B-Instr | 54.69 | 0.47 | 1.88 | 7.52 Mixtral-8x7B-Ins | 125.00 | 1.07 | 4.29 | 17.18 Qwen2.5-32B-Inst | 250.00 | 2.15 | 8.59 | 34.36 Qwen2.5-72B-Inst | 312.50 | 2.68 | 10.74 | 42.95 -----------------+------------+-------+--------+-------- KV sizes in GB; the per-token column is MiB per 1000 tokens. Table 3 converts the measured payload sizes into transfer times, which is the quantity a selection decision compares against the cost of recomputing or re-retrieving the same payload. Y. Mo Expires 23 April 2027 [Page 15] Internet-Draft Agent Service Characteristics October 2026 Payload | 1 Gbps | 10 Gbps | 40 Gbps | 100 Gbps -------------------+--------+---------+---------+--------- 0.5 GB, edge model | 6.7 | 0.7 | 0.2 | 0.1 5.6 GB, 7B 4-bit | 74.3 | 7.4 | 1.9 | 0.7 15.2 GB, 7B FP16 | 203.1 | 20.3 | 5.1 | 2.0 65.5 GB, 32B | 873.7 | 87.4 | 21.8 | 8.7 145.4 GB, 72B | 1938.8 | 193.9 | 48.5 | 19.4 688.6 GB, 671B | 9181.2 | 918.1 | 229.5 | 91.8 -------------------+--------+---------+---------+--------- Seconds, at 60% of the line rate. 6.2. Software Artifacts and Protocol Implementations For contrast, Table 4 gives the size of repositories that implement agent protocols, tool servers, and agent frameworks. Component | Repo size | Stars --------------------------------+-----------+------- modelcontextprotocol/python-sdk | 16.1 MB | 24.4k modelcontextprotocol/servers | 28.9 MB | 90.6k a2aproject/A2A | 33.8 MB | 25.9k openai/openai-agents-python | 62.4 MB | 29.7k microsoft/autogen | 145.2 MB | 61.2k crewAIInc/crewAI | 287.2 MB | 59.1k All-Hands-AI/OpenHands | 431.9 MB | 89.3k run-llama/llama_index | 486.7 MB | 52.3k langchain-ai/langchain | 588.9 MB | 147.2k --------------------------------+-----------+------- Repo sizes include history. Protocol implementations and tool servers are two to four orders of magnitude smaller than model artifacts. The storage and transfer burden of an agent service is therefore dominated by models and session state, while the shape of its traffic -- connection lifetime, message mix, directionality -- is determined by the protocol implementation in use. 6.3. Implications for CATS Four implications follow. * The quantities that a decision compares are of the same order of magnitude. Transferring a mid-size artifact at 10 Gbps takes on the order of tens of seconds (Table 3), and re-establishing the state that a transfer would avoid costs accelerator time of the same order. Neither dimension dominates by default, so a decision taken on network metrics alone or on computing metrics alone is equally likely to be wrong. Y. Mo Expires 23 April 2027 [Page 16] Internet-Draft Agent Service Characteristics October 2026 * Size awareness is unavoidable. A normalized score that does not know whether the payload is 0.9 GB or 689 GB cannot produce the same decision for both, and the same is true of the KV state, which ranges from about 1.6 GB to about 43 GB at a context of 128,000 tokens in the measured set. * State location changes within a session. Because the KV state grows with context, and because a long idle period can cause cached state to be evicted, affinity is not a property that can be fixed at session establishment; it must be re-evaluated, and the cost of re-establishing it must be accounted for. * The traffic caused by steering decisions is itself observable. Bulk state transfers and re-prefills triggered by a change of instance are attributable to that decision, which makes them measurable, and therefore manageable, provided that they are measured separately from background traffic. 7. Measurement Requirements A CATS system cannot improve what it does not measure. This section states what needs to be measured, where, and for what purpose. It defines no metric encoding and requests no registry entry; the intent is that the identified quantities can be carried as metrics of the existing framework [I-D.ietf-cats-metric-definition] and reported through the existing OAM functions [I-D.ietf-cats-oam-fw]. Measurement | Unit | Point and use ---------------------+-------------+------------------------- Model artifact load | bytes, s | C-SMA; storage metric State transfer: KV, | bytes, s, | C-SMA, C-NMA; session, index | by class | transfer cost, affinity Recompute and | count, s | C-SMA; transfer re-prefill events | | versus recompute Tool call round trip | s, by class | C-TC; forwarding and | | per-step budget Context size and | tokens, | C-SMA; storage metric growth rate | bytes | Session and turn | count, s | C-TC; budget, long tail lifecycle | | Selection changes | count, | C-PS; stability and cost | reason | Per-step latency | s, by phase | end to end, OAM; decomposition | | task objective ---------------------+-------------+------------------------- Y. Mo Expires 23 April 2027 [Page 17] Internet-Draft Agent Service Characteristics October 2026 The following properties are required for the measurement to be useful. * Traffic must be attributable. A measurement that cannot say which session, turn, step, or communication mode produced a byte cannot support a per-session objective, a budget, or a decision about affinity (R16). * Bulk transfers caused by steering decisions must be distinguishable from background traffic, otherwise the cost of a decision is invisible at the moment when the decision is made (R19, R21). * The tail must be measured, not only the mean. Because task size and idle duration are heavy-tailed (Section 5), objectives and reports that use averages hide the sessions that dominate cost and latency (R17). * Measurement must not become an information channel. State reuse is observable through timing, and multi-tenant measurements must be aggregated so that they do not disclose session content or identity. * Measurement must express the mode. A streaming agent-to-agent response, a person-facing interactive session, and a background tool call have different objectives, and a single aggregate for all three cannot support a per-mode objective (R18). 8. Mapping of Agent Service Characteristics to CATS Dimensions The steering decision for an agent step can be expressed as the selection among candidate service contact instances of the instance that best satisfies the requirements of the step. The information needed for that selection falls into three dimensions: forwarding, computing, and storage and state. 8.1. Forwarding Dimension The forwarding dimension describes the ability of the network to deliver the step and its dependencies within the required time and without loss of correctness. Relevant aspects include the latency and jitter of the candidate paths, available bandwidth, loss, the determinism that the underlay can provide, and the reachability of tools or data sources that the step will contact. Y. Mo Expires 23 April 2027 [Page 18] Internet-Draft Agent Service Characteristics October 2026 For agent services, two aspects deserve emphasis. First, the paths that matter are frequently not single paths but sets of paths, because a turn fans out to several destinations. Second, the traffic of a step can be tolerable to buffering or to reduction of precision in a way that ordinary best-effort traffic is not, a property that has been described as a semantic contract between the application and the network [I-D.li-cats-aisemantic-contract]. 8.2. Computing Dimension The computing dimension describes what an instance can execute and how much of that capacity is currently available. It has two distinct parts. The first part is capability: the models, adapters, quantization variants, runtimes, and accelerators that the instance supports, together with the context lengths and tool integrations that it can serve. Capability is the constraint described in the previous section. The second part is capacity and load: queueing condition, accelerator availability, memory headroom for a growing KV cache, and the recent performance of comparable steps. Capacity and load are quantities that can be traded against latency or cost. 8.3. Storage and State Dimension The storage and state dimension describes which parts of the working set an instance already holds, and what it costs to have them there. For each element of the working set -- session history, retrieved context, tool results, KV cache, model artifacts -- the following aspects are relevant: * availability: whether the element is present at the instance, and under which key; * residency tier: where the element is held, for example in accelerator memory, host memory, a node-local cache, or a shared store; * retrieval cost: the time and bandwidth required to make the element usable at the instance, or, alternatively, the cost of recomputing it from the request; * freshness and consistency: whether the element reflects the current Y. Mo Expires 23 April 2027 [Page 19] Internet-Draft Agent Service Characteristics October 2026 session state, and under which conditions it may be reused; * sharing scope: whether the element may be reused across sessions, tenants, or users. 8.4. Characteristic-to-Dimension Matrix The following matrix summarizes which dimension each characteristic primarily constrains. "C" denotes a constraint that must be satisfied, "Q" denotes a quantity that can be optimized. Characteristic Forwarding Computing Storage/State ------------------------------------------------------------ Session state Q - C + Q Multi-stage execution C + Q C + Q Q Discrete tool calls C + Q Q Q Long-horizon, long tail Q Q Q Context growth Q C + Q C + Q Capability tier - C Q Communication mode C + Q Q - State persistence Q Q C + Q Locality/governance C C C Cost and energy Q Q Q Task-level objectives Q Q Q 9. Requirements The following requirements are derived from the characteristics above. They are stated using the conventions of BCP 14 [RFC2119] [RFC8174], are requirements on the design of a CATS system that supports agent services, and are intended as input to the metric, data model, and selection work rather than as protocol requirements. 9.1. General Requirements R1. A CATS system SHOULD support selection at a granularity finer than the session, so that the requirements of individual steps or groups of steps can be taken into account. R2. A CATS system MUST be able to express constraints that cannot be traded against other quantities, in particular the capability required by a step and the locality constraints that apply to its state. R3. A CATS system SHOULD be able to attribute the traffic of a request to the session, turn, and step to which it belongs, to the extent needed to apply per-session budgets and to evaluate task-level objectives. Y. Mo Expires 23 April 2027 [Page 20] Internet-Draft Agent Service Characteristics October 2026 R4. A CATS system SHOULD avoid selection oscillation when the state of candidate instances changes on a timescale shorter than the instruction granularity of the selection function. 9.2. Forwarding Requirements R5. A CATS system SHOULD be able to consider a set of paths, rather than a single path, when a step fans out to multiple destinations. R6. A CATS system SHOULD be able to take into account the determinism and the isolation that candidate paths provide, when a step requires bounded latency rather than low average latency. R7. A CATS system SHOULD be able to distinguish traffic that tolerates buffering or precision reduction from traffic that does not, when such information is available to the network. 9.3. Computing Requirements R8. A CATS system MUST be able to treat the capability offered by an instance as a property that is matched against the requirement of a step, and not only as a quantity to be maximized. R9. A CATS system SHOULD expose and consume the load and availability of the specific resource that a step will consume, such as accelerator capacity or memory headroom for a growing KV cache, rather than only an aggregate utilization figure. R10. A CATS system SHOULD be able to account for the fact that the resources required by a step change as the step executes, for example when the memory used by a KV cache grows with the length of the conversation. 9.4. Storage and State Requirements R11. A CATS system SHOULD be able to determine, for a candidate instance, whether a given element of a session working set is available, and under which key or identifier it may be addressed. R12. A CATS system SHOULD be able to compare the cost of transferring a working set element to an instance with the cost of recomputing or re-retrieving it at that instance. R13. A CATS system MUST be able to enforce constraints on where a working set element may be placed, transferred, or reused, including jurisdiction, tenant, and device constraints. R14. A CATS system SHOULD be able to express state affinity, that is, the preference for an instance that already holds part of the working set, and to distinguish that preference from a hard locality constraint. Y. Mo Expires 23 April 2027 [Page 21] Internet-Draft Agent Service Characteristics October 2026 R15. A CATS system SHOULD be able to apply a ceiling on the cost or the energy attributed to a session or a turn, and to treat that ceiling as a constraint on the selection. R16. A CATS system SHOULD be able to attribute transfer, storage, and computing load to the session, turn, and step that produced it, to the extent needed to apply a per-session objective or budget. R17. A CATS system SHOULD evaluate its objectives at high percentiles as well as on average, because the distribution of task size, idle duration, and tool latency is heavy-tailed. R18. A CATS system SHOULD be able to distinguish the communication mode of a flow -- for example a persistent session transport, a request-response exchange, a streaming response, or an asynchronous callback -- to the extent needed to apply the objective that belongs to that mode. R19. A CATS system SHOULD be able to move session and model-side state in bulk, with integrity protection, bounded transfer time, and the ability to resume a transfer interrupted by a failure or by a change of the sender or receiver. R20. A CATS system SHOULD be able to treat context compression -- truncation, summarization, quantization of state, or replacement of history by retrieval -- as an alternative to transferring or recomputing state, and to account for the change that compression makes to the traffic and to the computing demand of a session. R21. A CATS system SHOULD measure the bulk transfers, recomputations, and re-prefills caused by its own selection decisions separately from background traffic, so that the cost of a decision is observable. R22. A CATS system SHOULD aggregate measurement data before exposing it, and SHOULD NOT require the content or the identity of a session in order to measure it. 10. Relationship to Existing and Ongoing Work This document complements, and does not replace, the following work. CATS framework, use cases, and metrics: The characteristics and requirements above are expressed using the identifiers, components, and metric levels of [I-D.ietf-cats-framework] and [I-D.ietf-cats-metric-definition]. No new functional component is introduced; the storage and state dimension is intended to be represented as an additional category of the existing metric framework. Y. Mo Expires 23 April 2027 [Page 22] Internet-Draft Agent Service Characteristics October 2026 Agent communication protocols: The communication modes described in Section 5, including MCP and A2A, are application-layer protocols defined elsewhere. This document does not define, profile, or modify them; it records the traffic patterns they produce so that a CATS system can be designed against them. Agent discovery: Discovery of agent resources, including the discovery of agents and of the information needed to reach them, is the subject of the proposed Discovery of Agents With Names (DAWN) working group [DAWN-Charter]. This document assumes that discovery has already happened, and concerns the selection among known instances as a function of forwarding, computing, and storage state. LLM inference requirements: The requirements that a large language model inference service places on the system and on the network have been analysed in [I-D.liu-nmrg-ai-llm-inference-requirements]. That analysis is complementary: it describes what the service needs, while this document describes what a CATS system needs in order to select an instance for the service, and neither replaces the other. Distribution mechanisms: Computing capability and metric information may also be distributed by other means and in other working groups. This document does not take a position on which mechanism carries such information, and considers that choice a matter for the working group. Agent identity and authorization: Identity, credentials, and the delegation of authority between agents are being addressed in other working groups. This document assumes that identity and authorization are handled by those mechanisms, and only requires that state access be subject to them. Other SDOs: Coordination of networking and computing is also being addressed outside the IETF. ITU-T SG13 consented draft new Recommendation ITU-T Y.3401 (ex YIMT2020-CNC-FW) on the coordination of networking and computing, ITU-T SG13 initiated a draft new Supplement ITU-T YsupCNC that provides a roadmap for that coordination, and ETSI ISG ISAC has approved Work Item 5. The CATS working group has exchanged liaison statements with both bodies [ITU-T-SG13-LS211] [ITU-T-SG13-LS214] [ETSI-ISAC-LS]. Those efforts address frameworks and coordination across networking and computing domains; this document addresses the information that a CATS system needs in order to select an instance for an agent service, and the two are complementary. Y. Mo Expires 23 April 2027 [Page 23] Internet-Draft Agent Service Characteristics October 2026 Individual documents on agent and inference traffic: Several individual documents describe parts of the problem addressed here, including agent-oriented token awareness [I-D.zhang-cats-token-aware-ts], the reference model for an AI-agent communication network [I-D.jiang-cats-reference-acn], distributed inference architectures [I-D.wang-cats-odsi] [I-D.li-cats-idn], and state distribution [I-D.li-cats-kv-cache-distribution]. This document consolidates the service characteristics that those documents assume. 11. Operational Considerations An operator that enables CATS for agent services needs to make choices that are not protocol questions. The granularity of selection, the triggers that cause a selection to be revisited, the weightings between the three dimensions, and the treatment of sessions whose state may not move are all deployment decisions. Two operational points are worth recording. First, the amount of state that is worth describing is bounded by the benefit of reusing it: an implementation is not expected to advertise every cache entry, and aggregation and summarization of state availability are expected to be necessary in large deployments. Second, the freshness of state and computing information determines what a selection function can usefully conclude; a decision made from stale information can be worse than a decision made from a static default, which argues for explicit consideration of freshness in the operational model [I-D.zhu-cats-metric-semantics]. 12. Security Considerations Enabling selection based on computing and storage state creates incentives to lie about that state. A participant that advertises capability or state that it does not have can attract traffic for sessions whose state it can then observe, and a participant that under-reports its load can attract traffic that it cannot serve. Claimed capability and claimed state availability therefore need the same kind of verification consideration as other advertised properties, and the mechanisms that verify them are a matter for the documents that define those advertisements rather than for this document. Describing sessions and steps to the network exposes information about users and applications: that a session exists, how long it is, which tools it uses, and how much state it holds. Least disclosure applies, and any information exchanged as a result of this document should be limited to what the selection decision requires. State handles share the properties of other identifiers: they may be guessed, replayed, or used as a capability. Possession of a state handle MUST NOT by itself authorize access to the state that it names. Y. Mo Expires 23 April 2027 [Page 24] Internet-Draft Agent Service Characteristics October 2026 13. Privacy Considerations The characteristics described in this document are derived from user behavior. Session length, step count, tool usage, context size, and locality constraints can reveal personal or business information even when the content of the session is not visible. Where such information is exposed to the network, it should be minimized, aggregated where possible, and subject to the same retention considerations as other operational data. 14. IANA Considerations This document has no IANA actions. 15. Normative References [I-D.ietf-cats-framework] Li, C., Du, Z., Boucadair, M., Contreras, L. M., et al., "A Framework for Computing-Aware Traffic Steering (CATS)", Work in Progress, Internet-Draft, draft-ietf-cats-framework-24, September 2026. [I-D.ietf-cats-metric-definition] Yao, K., et al., "CATS Metrics Definition", Work in Progress, Internet-Draft, draft-ietf-cats-metric-definition-11, September 2026. [I-D.ietf-cats-usecases-requirements] Yao, H., et al., "Computing-Aware Traffic Steering (CATS) Problem Statement, Use Cases, and Requirements", Work in Progress, draft-ietf-cats-usecases-requirements-14, September 2026. [RFC2119] Bradner, S., "Key words for use in RFCs to Indicate Requirement Levels", BCP 14, RFC 2119, DOI 10.17487/RFC2119, March 1997, . [RFC8174] Leiba, B., "Ambiguity of Uppercase vs Lowercase in RFC 2119 Key Words", BCP 14, RFC 8174, DOI 10.17487/RFC8174, May 2017, . 16. Informative References [I-D.ietf-cats-oam-fw] Xiong, Q., et al., "Computing-Aware Traffic Steering (CATS) Operations, Administration, and Maintenance (OAM) Framework", Work in Progress, Internet-Draft, draft-ietf-cats-oam-fw-01, July 2026. [DAWN-Charter] IETF, "Discovery of Agents With Names (DAWN) - Proposed Charter", 2026, . Y. Mo Expires 23 April 2027 [Page 25] Internet-Draft Agent Service Characteristics October 2026 [I-D.jiang-cats-reference-acn] Jiang, S., et al., "CATS Reference Model for AI-Agent Communication Network", Work in Progress, draft-jiang-cats-reference-acn-00, 2026. [I-D.li-cats-aisemantic-contract] Li, Q., et al., "Semantic-Driven Traffic Shaping Contract for AI Networks", Work in Progress, draft-li-cats-aisemantic-contract-01, August 2026. [I-D.li-cats-idn] Li, Z., et al., "A Framework of Intelligence Delivery Network (IDN) for Deep Learning Inference", Work in Progress, draft-li-cats-idn-01, August 2026. [I-D.li-cats-kv-cache-distribution] Li, Z., et al., "KV Cache Distribution for Distributed LLM Inference: Use Case and Requirements", Work in Progress, draft-li-cats-kv-cache-distribution-00, July 2026. [I-D.liu-nmrg-ai-llm-inference-requirements] Liu, Y., et al., "Requirements Analysis of System and Network for Large Language Model Inference Service", Work in Progress, Internet-Draft, draft-liu-nmrg-ai-llm-inference-requirements-02, 2026. [I-D.li-cats-task-segmentation-framework] Li, Q., et al., "A Task Segmentation Framework for Computing-Aware Traffic Steering", Work in Progress, draft-li-cats-task-segmentation-framework-02, July 2026. [I-D.ietf-cats-data-model] Lin, C., Yao, H., et al., "Data Model for Computing-Aware Traffic Steering (CATS)", Work in Progress, draft-ietf-cats-data-model-00, September 2026. [I-D.wang-cats-odsi] Wang, Y., et al., "An Architecture for Open, Decentralized, and Scalable Large Language Model Inference", Work in Progress, draft-wang-cats-odsi-01, August 2026. [I-D.zhang-cats-clients-request-packet] Zhang, Y., et al., "Carriage of CATS Service Identification and Request Constraints", Work in Progress, draft-zhang-cats-clients-request-packet-01, September 2026. [I-D.zhang-cats-token-aware-ts] Zhang, Y., et al., "A token-aware traffic steering solution for agent service", Work in Progress, draft-zhang-cats-token-aware-ts-00, July 2026. [I-D.zhu-cats-metric-semantics] Zhu, M., "Operational Semantics for CATS Metric Consumption", Work in Progress, draft-zhu-cats-metric-semantics-01, August 2026. [ETSI-ISAC-LS] ETSI ISG ISAC, "Liaison statement to IETF CATS on the approval of ETSI ISG ISAC Work Item 5", July 2025. Y. Mo Expires 23 April 2027 [Page 26] Internet-Draft Agent Service Characteristics October 2026 [ITU-T-SG13-LS211] ITU-T SG13, "Liaison statement SG13-LS211 to IETF CATS on the consent of draft new Recommendation ITU-T Y.3401 (ex YIMT2020-CNC-FW) on the coordination of networking and computing", August 2024. [ITU-T-SG13-LS214] ITU-T SG13, "Liaison statement SG13-LS214 to IETF CATS on the initiation of draft new Supplement ITU-T YsupCNC providing a roadmap for the coordination of networking and computing", August 2024. [MCP] Model Context Protocol, "Specification, revision 2025-06-18", June 2025, . [A2A] A2A Project, "Agent2Agent (A2A) Protocol Specification", . [MODELREG] "Model registry API (mirror)", model metadata and artifact sizes, measurements taken on 26 September 2026, . [GHAPI] "Code hosting API", repository sizes and popularity, measurements taken on 26 September 2026, . Acknowledgments The authors would like to thank the participants of the CATS working group for the discussions that shaped this document. Authors' Addresses Y. Mo Huazhong University of Science and Technology Email: moyj@hust.edu.cn Y. Mo Expires 23 April 2027 [Page 27]