Run Qwen 3.5 on your iPhone: the model behind Orvena
If you already run Qwen on a laptop, this is the same family of model in your pocket, wired to your calendar.
What Qwen 3.5 4B is
Qwen is a family of open-weight language models from Alibaba's Qwen team. "3.5" is the generation and "4B" is the size: about four billion parameters, which is the largest that runs comfortably in an iPhone's 8 GB of memory with room left for iOS. The model is multimodal, so it reads images as well as text. It can reason step by step before answering. And it was trained across a wide range of languages, which is part of how Orvena holds a conversation in thirty-one of them.
Why this model
Open weights. The weights are published, so they can be downloaded to the phone and run there. That is the precondition for everything else Orvena promises: a model cannot run on your device if you cannot have the weights.
It fits. At four bits per weight, the model is about 3.5 GB on disk and fits in memory alongside the app. Larger Qwen models exist and are better, but they do not fit in a phone today.
It uses tools. The model is trained to call functions, and function calling is what turns a chatbot into an assistant. When you ask Orvena to move a meeting, the model decides to call the calendar tool with the right arguments, reads the result, and tells you what happened. That is a skill the model has, not a trick the app performs around it.
How it runs on the phone
Orvena downloads the weights once, over Wi-Fi, from a pinned and checksummed source. It runs them with MLX, Apple's framework for machine learning on Apple silicon, on the iPhone's graphics processor. No network is needed afterwards; airplane mode is a fair test.
On an iPhone 16 Pro the model writes at roughly fifteen tokens a second, which is faster than most people read. The first words of a short answer appear within a second or two. Long conversations take longer to read in before the answer starts, and asking the model to think first adds a few seconds of reasoning.
What a 4B model does well, and where it stops
It is good at the things people actually ask an assistant to do on a phone: scheduling, drafting and rephrasing, summarising a page or a screenshot, reading a photo, following a multi-step request, switching languages mid-conversation. It is weaker than a data-center model at long essays and at obscure facts it never saw in training. For current information, ask it to search; the query goes to the search provider and the reading happens on the phone.
Compared with running Qwen on a Mac
On a Mac you choose the size, and the larger Qwen models are stronger. On a phone the 4B is the practical ceiling for now. What the phone adds is everything around the model: your calendar, reminders, photos, health summaries, maps, alarms, and a receipt for each action the model takes. Orvena is not a model runner with a chat box; it is an assistant built around one model and the phone it lives on.
Bigger models when you want them
If a request needs more than a phone-sized model can give, you can connect a cloud model with your own OpenRouter key, including larger Qwen models. Orvena asks for consent per provider, and again before sensitive history would leave the device. The default stays local.
Common questions
Which model does Orvena run on the phone?
Qwen 3.5 4B, quantized to four bits and run with MLX. Larger models are available optionally through your own OpenRouter key.
Does the model download every time?
No. The weights download once. After that the model runs with no network at all.
Is my data used to train Qwen?
No. The model runs on your phone. Nothing you type or say is transmitted to Alibaba, to Orvena, or to anyone else, so there is nothing for anyone to train on.
Orvena is free on the App Store for iPhone 15 Pro and newer. The model downloads once and runs on the phone.
Download on the App Store