How to make AI agent harness in Swift

Issue #1068

An AI agent is really just a loop. Your code sends a conversation to a language model, the model asks to run a tool, your code runs it and sends the result back, and this repeats until the model has enough to answer. The code around that loop, the request and response types, the tool dispatch, memory across turns, is what people mean by an AI harness.

There is no official Anthropic or OpenAI SDK for Swift, so building one for an iOS or macOS app means working directly on top of URLSession and Codable. It sounds like a lot of plumbing, but the shape is small once you see it: a message model, a client that can send one request, tool definitions, the loop itself, how to load a skill on demand, and a way to keep a long conversation from outgrowing its context window. That’s the order this article covers them in.

The message model

The Messages API works with a conversation: an array of messages, each with a role and a list of content blocks. A block can be plain text, a request to call a tool, or the result of a tool call. That variability is the first wrinkle in Swift, since a single array element type has to represent three different shapes:

swift
enum ContentBlock: Codable {
    case text(String)
    case toolUse(id: String, name: String, input: [String: JSONValue])
    case toolResult(toolUseId: String, content: String, isError: Bool)
}

struct Message: Codable {
    enum Role: String, Codable { case user, assistant }
    let role: Role
    let content: [ContentBlock]
}

Codable conformance for ContentBlock is mechanical once you write it: decode the type field, switch on it, decode the payload for that case, and mirror the same switch on encode. It’s boilerplate you write once and never look at again, so it’s left out here. JSONValue is a small helper enum, string, number, bool, object, array, null, that stands in for tool input and schemas, since those are arbitrary JSON that no fixed Swift type can describe. It’s a well known pattern, a few dozen lines, and you’ll find a reference implementation the first time you search for “Swift Codable arbitrary JSON.”

A client that sends one request

Before touching tools or loops, get a single round trip working. It is just a POST with three headers and a JSON body:

swift
struct MessageRequest: Encodable {
    let model: String
    let maxTokens: Int
    let system: String?
    let messages: [Message]
    let tools: [ToolDefinition]

    enum CodingKeys: String, CodingKey {
        case model, system, messages, tools
        case maxTokens = "max_tokens"
    }
}

struct MessageResponse: Decodable {
    let content: [ContentBlock]
    let stopReason: String?

    enum CodingKeys: String, CodingKey {
        case content
        case stopReason = "stop_reason"
    }
}

struct ClaudeClient {
    let apiKey: String
    private let session = URLSession.shared
    private let endpoint = URL(string: "https://api.anthropic.com/v1/messages")!

    func send(_ request: MessageRequest) async throws -> MessageResponse {
        var urlRequest = URLRequest(url: endpoint)
        urlRequest.httpMethod = "POST"
        urlRequest.setValue("application/json", forHTTPHeaderField: "Content-Type")
        urlRequest.setValue(apiKey, forHTTPHeaderField: "x-api-key")
        urlRequest.setValue("2023-06-01", forHTTPHeaderField: "anthropic-version")
        urlRequest.httpBody = try JSONEncoder().encode(request)

        let (data, response) = try await session.data(for: urlRequest)
        guard let http = response as? HTTPURLResponse, 200..<300 ~= http.statusCode else {
            throw ClaudeError.badStatus((response as? HTTPURLResponse)?.statusCode ?? -1, data)
        }
        return try JSONDecoder().decode(MessageResponse.self, from: data)
    }
}

Keep the API key out of source control: read it from the Keychain, not a literal string in a committed file. Run this once with an empty tools array and a single user message, print response.content, and you have proven the transport works. Everything from here is about making the loop smarter, not about the network call itself.

Tools

A tool is a name, a description, and a JSON Schema describing its arguments. The description matters more than it looks: the model decides whether to call a tool almost entirely from that text, so say plainly when it should be used, not just what it does.

Model the contract as a small protocol, separate from the wire format:

swift
protocol AgentTool {
    var name: String { get }
    var description: String { get }
    var inputSchema: JSONValue { get }
    func run(input: [String: JSONValue]) async throws -> String
}

struct ToolDefinition: Encodable {
    let name: String
    let description: String
    let inputSchema: JSONValue

    enum CodingKeys: String, CodingKey {
        case name, description
        case inputSchema = "input_schema"
    }
}

A harness is more convincing with two tools that behave nothing alike, so here is a web search tool that reaches out to the network, and a grep tool that searches the local filesystem instead. First, search:

swift
struct WebSearchTool: AgentTool {
    let name = "web_search"
    let description = "Search the web for current information. Call this when the question needs facts newer than your training or you are unsure."
    let inputSchema = JSONValue.object([
        "type": .string("object"),
        "properties": .object([
            "query": .object(["type": .string("string"), "description": .string("The search query")])
        ]),
        "required": .array([.string("query")])
    ])

    let search: SearchClient

    func run(input: [String: JSONValue]) async throws -> String {
        guard case .string(let query) = input["query"] else {
            throw ToolError.missingArgument("query")
        }
        let results = try await search.search(query: query)
        return results.map { "\($0.title) — \($0.snippet) (\($0.url))" }.joined(separator: "\n")
    }
}

SearchClient is just a protocol you implement against whatever search API you have, your own backend or a provider’s REST endpoint. The tool itself doesn’t care what is behind it, which is the point of keeping tools thin: they translate between the model’s request and whatever service already exists in your app.

Grep is the opposite kind of tool, no network at all, just shelling out to the grep binary that’s already on the machine:

swift
struct GrepTool: AgentTool {
    let name = "grep"
    let description = "Search files under a directory for lines matching a pattern. Call this to find where something is defined or used before answering."
    let inputSchema = JSONValue.object([
        "type": .string("object"),
        "properties": .object([
            "pattern": .object(["type": .string("string"), "description": .string("Pattern to search for")]),
            "path": .object(["type": .string("string"), "description": .string("Subdirectory to search, relative to the project root")])
        ]),
        "required": .array([.string("pattern")])
    ])

    let rootURL: URL

    func run(input: [String: JSONValue]) async throws -> String {
        guard case .string(let pattern) = input["pattern"] else {
            throw ToolError.missingArgument("pattern")
        }
        let subpath = if case .string(let value)? = input["path"] { value } else { "" }
        let searchRoot = rootURL.appendingPathComponent(subpath)

        let process = Process()
        process.executableURL = URL(fileURLWithPath: "/usr/bin/grep")
        process.arguments = ["-rn", "--include=*.swift", pattern, searchRoot.path]

        let output = Pipe()
        process.standardOutput = output
        process.standardError = Pipe()
        try process.run()
        process.waitUntilExit()

        let data = output.fileHandleForReading.readDataToEndOfFile()
        let text = String(data: data, encoding: .utf8) ?? ""
        return text.isEmpty ? "No matches" : text
    }
}

Process is synchronous, so waitUntilExit() blocks the calling thread until grep finishes, which is fine for a short-lived subprocess like this one. For a tool that could run long, you’d move it off to a background queue or wrap it with a continuation instead of blocking directly inside an async function.

A registry collects both tools, exposes their definitions for the request body, and dispatches a call by name:

swift
struct ToolRegistry {
    private var tools: [String: AgentTool] = [:]

    mutating func register(_ tool: AgentTool) {
        tools[tool.name] = tool
    }

    var definitions: [ToolDefinition] {
        tools.values.map { ToolDefinition(name: $0.name, description: $0.description, inputSchema: $0.inputSchema) }
    }

    func execute(name: String, input: [String: JSONValue]) async -> (output: String, isError: Bool) {
        guard let tool = tools[name] else {
            return ("No tool named \(name)", true)
        }
        do {
            return (try await tool.run(input: input), false)
        } catch {
            return ("\(error)", true)
        }
    }
}

Returning an error string with isError: true instead of throwing out of the loop matters. The model can read that a tool failed and try a different approach or ask the user for clarification, but only if the failure comes back as a normal tool result rather than crashing your app.

The loop

This is the part that turns a single API call into an agent. The model’s response carries a stop_reason. When it is tool_use, the response contains one or more blocks asking you to run a tool; you execute them, wrap the outputs as tool_result blocks, and send the whole conversation back. When it is end_turn, the model is done and you read the text out.

swift
final class Agent {
    private let client: ClaudeClient
    private let tools: ToolRegistry
    private let systemPrompt: String
    private var messages: [Message] = []

    init(client: ClaudeClient, tools: ToolRegistry, systemPrompt: String) {
        self.client = client
        self.tools = tools
        self.systemPrompt = systemPrompt
    }

    func send(_ text: String) async throws -> String {
        messages.append(Message(role: .user, content: [.text(text)]))

        while true {
            let request = MessageRequest(
                model: "claude-opus-5",
                maxTokens: 4096,
                system: systemPrompt,
                messages: messages,
                tools: tools.definitions
            )
            let response = try await client.send(request)
            messages.append(Message(role: .assistant, content: response.content))

            guard response.stopReason == "tool_use" else {
                return response.content.compactMap {
                    if case .text(let value) = $0 { return value }
                    return nil
                }.joined()
            }

            var results: [ContentBlock] = []
            for block in response.content {
                if case .toolUse(let id, let name, let input) = block {
                    let (output, isError) = await tools.execute(name: name, input: input)
                    results.append(.toolResult(toolUseId: id, content: output, isError: isError))
                }
            }
            messages.append(Message(role: .user, content: results))
        }
    }
}

Three details here are easy to get wrong and will silently break the conversation if you do. Always append the full response.content, not just the text you extracted from it; the tool_use blocks need to stay in history so the model can see what it already asked for. Every tool_result must carry the tool_use_id of the call it answers, since the model can ask for several tools in one turn and matches results back by id. And all the results for one turn go into a single user message, not one message per tool, or the model gradually stops making parallel tool calls.

For an app, Agent is the thing a view model holds onto: one instance per conversation, messages growing as the user and the model go back and forth, send(_:) called once per user turn.

Loading a skill

Some tasks need more than a tool call, they need a whole page of instructions: how your team writes release notes, the checklist for a code review, the format a report should follow. You could stuff all of that into the system prompt, but then every request pays for every skill’s tokens whether it’s relevant or not. A skill fixes this by staying small until it’s needed: only its name and a one-line description sit in context up front, and the full instructions load on demand.

Store each skill as a markdown file with a short frontmatter header and a body:

markdown
---
name: release-notes
description: Write App Store release notes from a list of git commits or bullet points.
---

Group changes under New, Improved, and More reliable. Write one benefit-first
sentence per bullet, starting with a verb. Never mention bug counts or version
numbers in the copy itself.

A library scans a folder of these on launch and keeps the parsed skills in memory:

swift
struct Skill {
    let name: String
    let description: String
    let instructions: String
}

struct SkillLibrary {
    private(set) var skills: [Skill] = []

    init(directory: URL) throws {
        let folders = try FileManager.default.contentsOfDirectory(at: directory, includingPropertiesForKeys: nil)
        for folder in folders {
            guard let raw = try? String(contentsOf: folder.appendingPathComponent("SKILL.md"), encoding: .utf8) else { continue }
            skills.append(Self.parse(raw))
        }
    }

    private static func parse(_ raw: String) -> Skill {
        let sections = raw.components(separatedBy: "---")
        let frontmatter = sections[1].split(separator: "\n")
        let body = sections[2...].joined(separator: "---").trimmingCharacters(in: .whitespacesAndNewlines)
        let name = frontmatter.first { $0.hasPrefix("name:") }?.dropFirst(5).trimmingCharacters(in: .whitespaces) ?? ""
        let description = frontmatter.first { $0.hasPrefix("description:") }?.dropFirst(12).trimmingCharacters(in: .whitespaces) ?? ""
        return Skill(name: name, description: description, instructions: body)
    }

    var summaries: String {
        skills.map { "- \($0.name): \($0.description)" }.joined(separator: "\n")
    }
}

Loading a skill turns out to be nothing new: it’s a tool, the same shape as WebSearchTool or GrepTool, except its description already lists every skill by name so the model knows what’s available without a separate lookup step:

swift
struct LoadSkillTool: AgentTool {
    let library: SkillLibrary

    var name: String { "load_skill" }
    var description: String {
        "Load the full instructions for a skill before starting a task it covers. Available skills:\n\(library.summaries)"
    }
    var inputSchema: JSONValue {
        .object([
            "type": .string("object"),
            "properties": .object([
                "name": .object(["type": .string("string"), "description": .string("The skill's name, exactly as listed")])
            ]),
            "required": .array([.string("name")])
        ])
    }

    func run(input: [String: JSONValue]) async throws -> String {
        guard case .string(let skillName) = input["name"] else {
            throw ToolError.missingArgument("name")
        }
        guard let skill = library.skills.first(where: { $0.name == skillName }) else {
            return "No skill named \(skillName)"
        }
        return skill.instructions
    }
}

Register it like any other tool, and the loop from the previous section handles the rest without changes: the model calls load_skill, the instructions come back as a tool_result, and from that point on they sit in the conversation exactly as if you had pasted them into the prompt yourself. A library of fifty skills costs a few lines of description per request, not fifty pages of instructions.

Context and memory

The loop above has a problem it doesn’t show: messages only grows. Every tool call and every result stays in the array forever, and every one of those tokens gets resent on every following request. Left alone, a long conversation eventually gets slow, expensive, and then simply too large for the model’s context window. A harness needs two separate mechanisms for this, one for the current conversation and one for what should survive past it.

The first is compaction: when the transcript passes some size, summarize the older turns into a few sentences and keep only the summary plus the most recent exchanges. It’s worth spending a cheaper, faster model on this, since a summary doesn’t need your main model’s reasoning:

swift
extension Agent {
    private static let maxTurnsBeforeCompaction = 20
    private static let turnsToKeepVerbatim = 6

    private func compactIfNeeded() async throws {
        guard messages.count > Self.maxTurnsBeforeCompaction else { return }

        let older = messages.prefix(messages.count - Self.turnsToKeepVerbatim)
        let recent = Array(messages.suffix(Self.turnsToKeepVerbatim))

        let summaryRequest = MessageRequest(
            model: "claude-haiku-4-5",
            maxTokens: 512,
            system: "Summarize this conversation in a few sentences. Keep any facts, decisions, or open questions.",
            messages: Array(older),
            tools: []
        )
        let summary = try await client.send(summaryRequest)
        let summaryText = summary.content.compactMap {
            if case .text(let value) = $0 { return value }
            return nil
        }.joined()

        messages = [Message(role: .user, content: [.text("Earlier conversation, summarized: \(summaryText)")])] + recent
    }
}

Call compactIfNeeded() at the top of send(_:), before building the request. This keeps the conversation coherent without keeping every tool call verbatim.

The second mechanism is memory that outlives a single conversation entirely, the kind of thing a user shouldn’t have to repeat every time they open the app. That’s just a small file on disk, read into the system prompt at the start of a session and appended to when something worth remembering comes up:

swift
struct MemoryStore {
    private let url: URL

    func notes() -> String {
        (try? String(contentsOf: url, encoding: .utf8)) ?? ""
    }

    func remember(_ note: String) {
        let updated = notes() + "\n- \(note)"
        try? updated.write(to: url, atomically: true, encoding: .utf8)
    }
}

Build the system prompt as the fixed instructions plus whatever notes exist, systemPrompt + "\n\nWhat you know about this user:\n" + memory.notes(). You can write to it yourself from app logic, or go one step further and give the model a remember tool that calls memory.remember(_:), so it decides on its own what’s worth keeping between sessions rather than you guessing in advance.

Between compaction and memory, the harness ends up with three tiers of state instead of one: the live transcript for the current exchange, a rolling summary for the conversation so far, and a small durable file for facts that matter beyond any one conversation. That separation is what keeps an agent usable past the first few dozen turns, and it’s the part that’s easy to skip when you’re only testing with a handful of messages.

Written by

I’m open source contributor, writer, speaker and product maker.

Start the conversation