How to use Foundation Models in iOS

Issue #1063

Adding AI to an app used to mean calling a cloud API and accepting the latency, cost, and privacy trade-offs that come with it. The Foundation Models framework, introduced in iOS 26, removes that trade-off by giving you direct Swift access to the on-device model behind Apple Intelligence, no network call and no per-token bill required.

This article covers the framework end to end: checking availability, running a session, shaping output into Swift types, streaming, tool calling, and performance tuning, plus what iOS 27 added on top.

What you’re actually calling

SystemLanguageModel is Apple’s on-device model, exposed through the FoundationModels framework. It’s a small model compared to something like GPT-4, tuned for on-device tasks such as summarizing, rewriting, classifying, extracting structured data, and answering questions grounded in content you provide. It is not a general knowledge engine, and Apple doesn’t pretend otherwise: don’t expect it to know yesterday’s news or recite obscure trivia. What it’s good at is taking text you hand it and transforming that text into a shape your app can use, quickly and privately.

Because the model lives on the device, it only works where Apple Intelligence works. That means a specific list of supported hardware, the feature toggled on in Settings, enough battery, and the device not sitting in Game Mode. The model itself can also take a moment to download the first time Apple Intelligence is enabled. All of this means your very first job, before writing any generation code, is to ask the model whether it’s ready.

import FoundationModels

let model = SystemLanguageModel.default

switch model.availability {
case .available:
    // proceed
    break
case .unavailable(.deviceNotEligible):
    print("This device doesn't support Apple Intelligence.")
case .unavailable(.appleIntelligenceNotEnabled):
    print("Turn on Apple Intelligence in Settings to use this feature.")
case .unavailable(.modelNotReady):
    print("The model is still downloading. Try again shortly.")
case .unavailable:
    print("The model is unavailable right now.")
}

Treat this switch as mandatory, not defensive boilerplate. A real share of your users will hit one of the unavailable branches, and a feature that silently does nothing is worse than one that explains itself.

Starting a session

Once the model is available, everything runs through a LanguageModelSession. A session holds conversation state, so if you want follow-up questions to understand earlier context, keep the same session alive rather than creating a new one per request.

@State private var session = LanguageModelSession()

func askFollowUp(_ text: String) async throws -> String {
    let response = try await session.respond(to: text)
    return response.content
}

Most sessions benefit from instructions: a block of text that sets the model’s role, tone, and rules for every prompt in that session, similar to a system prompt in other LLM APIs.

let session = LanguageModelSession(instructions: """
    You are a cooking assistant. You suggest recipes based on
    ingredients the user already has. Keep suggestions realistic
    for a home kitchen and never invent ingredients the user
    didn't mention.
    """)

Instructions are trusted differently than user input: the model treats them as the developer’s rules, which is exactly why they’re the right place to put constraints you don’t want a clever prompt to override. LanguageModelSession also accepts a tools array and a guardrails setting, both of which I’ll come back to. Guardrails can’t be turned off entirely, since Apple filters unsafe content by design, but you can relax them with .permissiveContentTransformations for tasks like summarizing or rewriting text the user already wrote and already trusts.

Getting structured output instead of a paragraph

Plain text is fine for a chat bubble, but most real features need shaped data: a title, a list of steps, a price, a category. Rather than parsing that out of prose yourself, or fighting a model into emitting valid JSON, Foundation Models lets you describe a Swift type and generate it directly.

@Generable
struct Recipe {
    @Guide(description: "A short, appetizing name for the dish.")
    let title: String

    @Guide(description: "Ingredients with quantities, one per line.")
    @Guide(.count(3...8))
    let ingredients: [String]

    @Guide(description: "Numbered cooking steps in order.")
    let steps: [String]
}

let response = try await session.respond(
    to: "Suggest a dinner using chicken, rice, and broccoli.",
    generating: Recipe.self
)

let recipe = response.content
print(recipe.title)

@Generable builds the schema the model uses to constrain its own output, and @Guide attaches hints and constraints to individual properties, things like a description, a count range, or a numeric bound. The result is a real Recipe value, not a string you have to hope parses correctly. This works for enums too, which is useful whenever you want the model to pick from a fixed set of categories rather than free-form text.

@Generable
enum MealType: String, CaseIterable {
    case breakfast
    case lunch
    case dinner
    case snack
}

Nesting @Generable types inside each other lets you model something as rich as a full itinerary or a multi-course menu, and the framework will populate the whole tree in one generation call.

Shaping the prompt itself

Once your instructions and output type are set, the wording of the prompt still matters. Prompt, built with a result builder, lets you compose a prompt out of conditional pieces the same way you’d compose a SwiftUI view.

let vegetarian = true

let prompt = Prompt {
    "Suggest a dinner using chicken, rice, and broccoli."
    if vegetarian {
        "The user is vegetarian, so replace the chicken with a plant-based protein."
    }
}

For trickier tasks, one-shot or few-shot prompting, showing the model a fully worked example alongside the request, noticeably improves consistency. You can drop a @Generable value straight into the prompt builder as that example.

let prompt = Prompt {
    "Suggest a dinner using chicken, rice, and broccoli."
    "Match this style and level of detail, but invent new content:"
    Recipe.example
}

This costs a few more tokens per request, but for anything user-facing where consistency matters more than raw speed, it’s usually worth it.

Streaming as the model writes

A generation request can take a few seconds, and staring at a blank screen for that long feels broken even when it’s working correctly. streamResponse gives you the same generation, but delivered incrementally as PartiallyGenerated values you can render as they arrive.

let stream = session.streamResponse(to: prompt, generating: Recipe.self)

for try await partial in stream {
    self.recipe = partial.content
}

Every property on a partially generated type is optional, since the model may not have produced it yet, so your view code unwraps as it goes.

if let title = recipe?.title {
    Text(title)
}
if let ingredients = recipe?.ingredients {
    ForEach(ingredients, id: \.self) { Text($0) }
}

Pairing this with .contentTransition(.opacity) and a light .animation on the container gives you the same word-by-word reveal you’d get from a chat app, without writing any manual diffing.

Giving the model tools

The base model only knows what’s in its training data and whatever you put in the prompt. When a feature needs live or private information, your app’s HealthKit data, a database lookup, the current weather, you give the model a Tool. The model itself decides when calling that tool is necessary to answer the prompt.

import FoundationModels

final class PantryLookupTool: Tool {
    let name = "pantryLookup"
    let description = "Returns the ingredients currently in the user's pantry."

    @Generable
    struct Arguments {}

    func call(arguments: Arguments) async throws -> String {
        let items = await PantryStore.shared.currentItems()
        return "Pantry contains: \(items.joined(separator: ", "))"
    }
}

let session = LanguageModelSession(
    tools: [PantryLookupTool()],
    instructions: "Suggest recipes using only what's in the user's pantry."
)

name and description aren’t cosmetic, they’re the only information the model has about when and why to reach for the tool, so write the description the way you’d write documentation for another developer. Arguments follow the same @Generable pattern as structured output, so a tool that needs input from the model, like a search term or a date range, declares it as a @Guided property on its Arguments struct. Errors thrown from call(arguments:) surface back through LanguageModelSession.ToolCallError, so you can catch tool-specific failures separately from general generation errors.

catch let error as LanguageModelSession.ToolCallError {
    print("Tool \(error.tool.name) failed: \(error.underlyingError)")
}

Performance habits worth adopting

Two small changes noticeably improve how a Foundation Models feature feels in practice. The first is prewarming: call session.prewarm() as soon as you know a generation is likely, for example when a user opens a screen with an “Ask” button, rather than waiting until they tap it. This loads the model into memory ahead of time and shaves real latency off the first response.

.task {
    session.prewarm()
}

The second is watching what you put in the prompt versus what you put in @Guide descriptions. Every word in the prompt costs tokens and processing time, and once your @Generable type already encodes the shape you want, you don’t need to repeat that structure in prose. Trust the schema to do the constraining and keep the prompt itself focused on the actual request.

GenerationOptions gives you a few more levers per request: temperature for how conservative or creative the output should be, maximumResponseTokens to cap length, and sampling, where .greedy sampling produces more deterministic, repeatable output, which is often what you want once a feature is tool-driven and structured rather than purely conversational.

let response = try await session.respond(
    to: prompt,
    generating: Recipe.self,
    options: GenerationOptions(sampling: .greedy)
)

What iOS 27 added

The iOS 26 version of the framework left tool calling entirely up to the model: you attached tools to a session, and from that point on the model alone decided whether, and how often, to call them. iOS 27 gives you a dial on that behavior. GenerationOptions.ToolCallingMode is a new per-request setting that lets you steer how aggressively the model reaches for tools on a given request, and the framework can automatically shift out of tool-calling mode after the first call so a request doesn’t loop indefinitely.

var options = GenerationOptions()
options.toolCallingMode = .required // steer this request's tool behavior explicitly

let response = try await session.respond(
    to: "Find three pantry-friendly dinners for tonight.",
    options: options
)

Because this lives on GenerationOptions rather than on the session, you can be strict about tool use on one prompt and let the model answer freely from context on the next, inside the same conversation.

iOS 27 also ships two ready-made tools from the Vision framework, so you no longer have to write your own recognition code for two extremely common cases. OCRTool reads text out of an image and hands back a string; BarcodeReaderTool scans machine-readable codes and returns each one’s decoded content along with its symbology.

import FoundationModels
import Vision

let session = LanguageModelSession(tools: [OCRTool(), BarcodeReaderTool()])

let response = try await session.respond(
    to: "Read this receipt photo and tell me the total and the date."
)

You attach them exactly like a tool you wrote yourself, and you can override each one’s name and description to bias the model toward or away from using it for a given app’s prompts.

Written by

I’m open source contributor, writer, speaker and product maker.

Start the conversation