Apple's on-device Foundation Model: what it can and can't do, from shipping it in AutoAlign

By Osama Mazhar · Published · 9 min read

I released AutoAlign in August 2025 to fix keystone distortion, the leaning walls and buildings in photos. Correcting perspective is most of what document scanning needs, so AutoAlign 4.0 added scanning. Then, to make the scanned documents useful, I added Ask Documents: a chat that answers questions about your documents using only the on-device model of Apple Intelligence, through Apple's Foundation Models framework.

This article covers what the model did well, where it fell short, and what I had to build around it. It's written for developers considering the framework, and for anyone wondering how capable a phone-sized language model really is.

Quick answer

Apple's on-device Foundation Model is good at short, structured jobs: filing a document into a fixed set of fields, picking which document a question is about, and answering a single question from a few passages it is given. It is weak at arithmetic, at long multi-turn conversations, and at anything that needs more than its roughly 4,096-token context window. In AutoAlign it usually finds the right document and answers within seconds. Follow-up questions are where it still struggles.

Key facts

ModelApple Intelligence on-device model, via SystemLanguageModel.default in the Foundation Models framework
Cloud useNone. AutoAlign never uses Private Cloud Compute.
Context windowAbout 4,096 tokens for instructions, prompt, image and answer combined
DevicesApple Intelligence devices, such as iPhone 15 Pro and later, on iOS 26 or later
ImagesOn iOS 27, where the model reports the vision capability, it can also look at the page
OutputTyped Swift structs through @Generable and @Guide, decoded greedily
RetrievalBM25 over NLTagger lemmas, fused with NLContextualEmbedding similarity by reciprocal-rank fusion
Works wellFinding the right document and answering one question about it
Works poorlyMulti-turn conversations, arithmetic, long inputs

What is Apple's on-device Foundation Model?

It is the compact language model that powers Apple Intelligence, opened to apps in iOS 26 through the Foundation Models framework. It runs on the device, needs no internet once downloaded, and costs nothing per request. An app opens a LanguageModelSession, gives it instructions, and asks it to respond, either in free text or in a structure the app defines.

Because it runs on the phone, documents never leave it. For a document app, that settled the choice before quality came into it.

What is the on-device model good at?

Structured answers with @Generable

Guided generation is the framework's best feature. You describe the shape of the answer as a Swift type, and the model is constrained to fill it. AutoAlign's answer type looks like this:

@Generable
struct ChatAnswer {
    @Guide(description: "The answer itself, in full sentences, in the language of the user's last message, its first words the fact asked for")
    var text: String
    @Guide(description: "False when the documents do not hold the answer")
    var found: Bool
    @Guide(description: "Titles of the documents the answer was taken from, exactly as they are listed", .maximumCount(3))
    var documents: [String]
}

The found flag lets the app say plainly when the answer isn't in your documents, without parsing "not found" in seven languages. The documents field is how every answer shows its source.

Deciding which document a message is about

Before answering, a separate small session reads each message with the list of your documents and works out which one it means ("my last receipt", "this one", "the slides") and whether you're asking, telling it something to keep, or renaming a document. Splitting this out worked better than asking the answering session to do it: when one session did both, it labelled nearly every question as a note to keep.

Answering one question from a few passages

Given the right passages, the model answers single factual questions well: a receipt total, a due date, a number on a card. This is the case that works in most of our tests, usually within seconds.

What is the on-device model bad at?

Multi-turn conversation

This is the biggest limit. The instructions, the document list, the retrieved passages and room for the answer take most of the 4,096 tokens. Earlier turns have to go. AutoAlign keeps the last two exchanges, then one, then none, until the next turn fits. A chain of follow-ups like "and the one before that?" or "compare it with last month" loses track.

Two things I learned the hard way:

Arithmetic

It cannot reliably add up a column of numbers. AutoAlign computes spending totals in code and puts them in the instructions, and the model is told not to add amounts up itself.

Long inputs

You cannot hand it all your documents. Search has to happen before the model sees anything, which is why retrieval matters more than the prompt.

How do you fit a document collection into 4,096 tokens?

By searching first, on the device, with Apple's own language tools:

  1. Text recognition. Vision's RecognizeDocumentsRequest reads each page, with its lines and layout.
  2. Keyword ranking. BM25 over words reduced to their dictionary form with NLTagger, so "paid" also finds "pay". Names found in the question count double.
  3. Meaning ranking. NLContextualEmbedding turns passages and the question into vectors, so "groceries" can find a supermarket receipt that never uses the word.
  4. Fusion. The two rankings are merged with reciprocal-rank fusion, which needs no weights to tune.

The top four passages go to the model, with the page itself as an image where the model can see. For budgeting I estimate about three characters per token and count an image as about 200 tokens.

What prompt mistakes does a small model make?

Each of these showed up in testing and needed a rule in the instructions:

How should an app check if the model is available?

Offer the feature only when it can answer now. AutoAlign shows the Ask button only when both are true:

SystemLanguageModel.default.availability == .available
    && SystemLanguageModel.default.supportsLocale()

Otherwise there's no button and no error screen, and Settings explains why: the device isn't eligible, Apple Intelligence is off, the model is still downloading, or the language isn't supported. Arabic, one of AutoAlign's seven languages, isn't supported by the model yet. The system keeps the list of supported languages, so it grows without an app update.

What's next?

Ask Documents is still experimental. The open problem is follow-up questions: keeping the thread of a conversation inside 4,096 tokens without the model repeating itself. Rolling summaries and structured memory of what was discussed are the next things to try.

If you have an iPhone 15 Pro or later, you can try it for free: free users can scan or import 4 documents and ask about them. I'd like to hear where it answers well and where it gets things wrong. Send feedback.

FAQ

How big is the context window of Apple's on-device Foundation Model?

About 4,096 tokens, shared by the instructions, the prompt, any image and the answer. In AutoAlign's document chat, instructions, retrieved passages and the reply leave little room for earlier turns.

Can Apple's on-device model answer questions about many documents?

Not by reading them all. The app has to search first and pass only the best few passages. AutoAlign uses BM25 keyword search fused with NLContextualEmbedding similarity, then gives the top four passages to the model. With that, it usually finds the right document and answers within seconds.

Is Apple's on-device model good at multi-turn conversation?

Not yet, in our experience. Single questions work well. A long chain of follow-ups loses track, because older turns have to be dropped to fit the 4,096-token window.

Does the Foundation Models framework send data to the cloud?

SystemLanguageModel runs on the device. Apple also offers larger models through Private Cloud Compute, but an app chooses which to use. AutoAlign uses only the on-device model, so documents and questions never leave the iPhone.

Which devices can run Apple's on-device Foundation Model?

Devices that support Apple Intelligence, such as iPhone 15 Pro and later, running iOS 26 or later with Apple Intelligence turned on and its model downloaded. Apps should check SystemLanguageModel.default.availability and supportsLocale() before offering the feature.

Can the on-device model do arithmetic?

Not reliably. AutoAlign computes spending totals in code and gives the model the result, and the instructions tell the model not to add amounts up itself.

Next reads

Try Ask Documents Free for 4 documents. iPhone 15 Pro or later.
Download