How Auto Align answers questions about your documents with on-device Apple Intelligence
Auto Align 4.0 does something new for a photo app: it keeps your documents. Scan a receipt, a letter or a set of slides, and the pages are cropped, straightened and filed. Then you can ask about them in plain words, like "what was the total on the Hofer receipt?" or "when is the insurance letter from?", and get an answer with the document it came from.
All of it runs on your iPhone. There is no Auto Align server and no account, and the app never uses Apple's cloud model. This article explains how that works, using Apple's Foundation Models framework, the on-device language model behind Apple Intelligence.
Auto Align reads each page's text on the device, then asks Apple's on-device model to name and file the document in a fixed format. When you ask a question, the app first finds the few passages that match it, then gives only those to the model to answer from. Nothing is sent anywhere. Ask Documents appears on devices that support Apple Intelligence, running iOS 26 or later, with Apple Intelligence turned on.
What is Apple's Foundation Models framework?
With iOS 26, Apple opened the language model that powers Apple Intelligence to apps. It is a compact model that lives on the device and runs on the iPhone's Neural Engine and GPU. An app talks to it through the Foundation Models framework: it opens a session, gives it instructions, and asks it to respond.
Three things make it a good fit for documents:
- It is private by design. The model runs locally, so text never has to leave the phone. Apple also offers larger models in Private Cloud Compute, but Auto Align deliberately uses only the on-device one.
- It needs no connection. Once Apple Intelligence has downloaded its model, answers are generated on the device, without the internet.
- It can answer in a structure, not just prose. The framework's guided generation lets an app describe the exact shape of the answer it wants, and the model is constrained to fill in that shape. No parsing of free text, no malformed output.
Step 1: reading the page
A language model cannot help with a document it cannot read, so the first step is text recognition. Auto Align uses Apple's Vision framework, which on iOS 26 recognises whole documents, including their lines and layout, on the device. The recognised text is stored with the page, so it can be searched later.
Step 2: filing the document
Once the text is read, Auto Align asks the on-device model to fill in a small document card:
- a short, human title of at most 60 characters
- a category: receipts, invoices and bills, cards, notes and whiteboards, slides and screens, letters and forms, IDs, or other
- who issued it, if it says
- the document's own date, if one appears on it
- up to three tags
Because of guided generation, the model cannot invent a ninth category or return a date in the wrong format. The app also asks for the model's most likely answer every time, not a creative one, so the same page is filed the same way twice. On devices where the model can also look at images, it is shown the first page as well as its text.
The model is not always available: the device may not support it, Apple Intelligence may be off, or the language may not be one it speaks. So filing has two fallbacks. The first matches the page against example documents using Apple's on-device sentence embeddings, and uses Apple's language tools to pick out names and dates. The second, which covers all seven of the app's languages, uses keyword lists and the shape of the page: a wide page with few words is probably a slide. Every device gets its documents filed; the model just does it better.
Step 3: finding the right passages
The on-device model has a small context window of about 4,096 tokens, roughly a few pages of text. You cannot hand it every document you own and ask a question. So before the model sees anything, Auto Align searches your documents itself, using two methods at once:
- Word matching. A classic ranking method (BM25) over the words in each page, reduced to their dictionary form, so "paid" also finds "pay". Names in your question, like a shop or a company, count for more.
- Meaning matching. Apple's contextual embeddings turn each passage and your question into vectors, so "how much did I spend on groceries" can find a supermarket receipt that never uses the word "groceries".
The two rankings are combined, and only the best few passages go to the model, together with a short list of your documents and the last few messages of the conversation. That is how a small model answers questions about a large collection.
Step 4: understanding what you mean
Before answering, the model reads your message once to work out two things: which documents it is about ("this one", "my last receipt", "the slides") and what you want. You might be asking a question, telling it something to remember ("this was the office lunch"), or renaming a document ("call it Car insurance 2026"). Notes are saved on the document itself, so they help both search and later answers.
Step 5: the answer
The answer also comes back in a fixed shape: the reply itself, whether the documents actually held the answer, and which documents it was taken from. That is why Auto Align can show the source document under each answer, and say plainly when the answer is not in your documents instead of guessing.
A few rules keep answers honest:
- No arithmetic by the model. Language models are poor at adding columns of numbers, so the app works out spending totals itself from the receipts and bills and gives the model the result.
- Only what is on the page. The model is told never to invent numbers or names, and to answer from the passages it was given.
- Your language. The answer comes back in the language you asked in, even when the document is in another one, with currencies kept.
Conversations are kept on the device, so you can pick one up again later.
Why does the Ask button only appear on some iPhones?
Ask Documents needs the on-device model to be ready. Auto Align shows the Ask button only when all of these are true:
- the device supports Apple Intelligence (for example iPhone 15 Pro and later, or iPads and Macs with an M-series chip)
- it runs iOS 26 or later
- Apple Intelligence is turned on in Settings and has finished downloading its model
- your language is one the model supports (Arabic is not, yet)
If any of these is missing, the button is simply not there. There is no screen that tells you no. Everything else, including scanning, filing, search and Spotlight, works the same on every device. Turn Apple Intelligence on, come back to the app, and the button appears.
What it cannot do
- It is labelled experimental for a reason: answers can be wrong. Check important figures against the document, which is always one tap away.
- It works best in English. Other supported languages work, with more mistakes.
- It only knows what is in your scanned documents. It does not search the web.
- It does not sync documents between your devices.
FAQ
Does Ask Documents send my documents to the cloud?
No. Auto Align uses only Apple's on-device model through the Foundation Models framework. It never uses Private Cloud Compute, and orbitaar has no server. Your documents, questions and answers stay on your iPhone.
Which devices support Ask Documents?
Devices that support Apple Intelligence, running iOS 26 or later with Apple Intelligence turned on, such as iPhone 15 Pro and later. On other devices the Ask button is not shown, and scanning, filing and searching work as usual.
Why don't I see the Ask button?
It appears only when the on-device model is ready: Apple Intelligence must be supported, turned on and finished downloading, and your language must be one the model supports. Arabic is not supported yet.
Is scanning still useful without Apple Intelligence?
Yes. Without the model, Auto Align files documents with Apple's on-device language tools and its own keyword rules, which cover all seven of the app's languages.
Next reads
- What is Auto Align? — the full app overview
- How do you fix leaning buildings in iPhone photos?
- Auto Align privacy policy