> Instead of forcing the entire model into DRAM, the full model is stored in flash memory (NAND). Because NAND-to-DRAM bandwidth is too slow to swap weights token by token, as standard MoE models require, AFM 3 Core Advanced makes routing decisions per prompt. A lightweight, dense block selects a fixed set of experts during initial processing, periodically reselecting them during generation. To minimize data movement, the model relies on a high percentage of always-active “shared experts” alongside input-dependent “routed experts” swapped into DRAM only when needed.
This is an interesting hybrid between MoE and managing entirely separate domain-specific models. Select the experts once, bring them into memory, and run inference for some period of time before re-evaluating. Saves having all experts in memory, but it's better than just selecting a whole model per query since you have a high number of small opaque experts that overlap and combine in interesting ways.
There is a probably a massive quality hit to doing this but it's interesting because it allows infinite scaling of model size.
"We do not use our users’ private personal data or user interactions when training our foundation models. We also respect the rights of web publishers to opt out of foundation model training."
What Apple consumers want from Apple is a premium product without all the enshittification found everywhere else. I don't know what Apple thinks it's doing by introducing so many ads into everything (maps now!) and shuffling their feet backwards on all their privacy positions.
In a world where every company steals all my data equally, and is ridden with the same crappy ads, why on Earth would I pay a hefty premium to Apple? It's just so stupid and short-sighted in terms of product differentiation.
So far Apple has been using default-checked checkbox during onboarding process for OS-level “help us improve…” diagnostics collection. Not sure about this specifically.
Regarding the architecture:
> Instead of forcing the entire model into DRAM, the full model is stored in flash memory (NAND). Because NAND-to-DRAM bandwidth is too slow to swap weights token by token, as standard MoE models require, AFM 3 Core Advanced makes routing decisions per prompt. A lightweight, dense block selects a fixed set of experts during initial processing, periodically reselecting them during generation. To minimize data movement, the model relies on a high percentage of always-active “shared experts” alongside input-dependent “routed experts” swapped into DRAM only when needed.
This is an interesting hybrid between MoE and managing entirely separate domain-specific models. Select the experts once, bring them into memory, and run inference for some period of time before re-evaluating. Saves having all experts in memory, but it's better than just selecting a whole model per query since you have a high number of small opaque experts that overlap and combine in interesting ways.
There is a probably a massive quality hit to doing this but it's interesting because it allows infinite scaling of model size.
Changed again?
"We do not use our users’ private personal data or user interactions when training our foundation models. We also respect the rights of web publishers to opt out of foundation model training."
Further down. Point 4 under Responsible AI
Your private personal data and interactions are never used to train our foundation models unless you explicitly choose to help improve them.
That's practically the opposite of what your editorialized submission title says.
Title seems misleading if not outright wrong, having read the article.
what a coincidence, i thought we were going to slow down training (amodei, altman et al.)
Apple keeps taking potshots at their own feet.
What Apple consumers want from Apple is a premium product without all the enshittification found everywhere else. I don't know what Apple thinks it's doing by introducing so many ads into everything (maps now!) and shuffling their feet backwards on all their privacy positions.
In a world where every company steals all my data equally, and is ridden with the same crappy ads, why on Earth would I pay a hefty premium to Apple? It's just so stupid and short-sighted in terms of product differentiation.
They changed
> We do not use our users’ private personal data or user interactions when training our foundation models.
https://web.archive.org/web/20260829051311/https://machinele...
to
> Your private personal data and interactions are never used to train our foundation models unless you explicitly choose to help improve them.
https://machinelearning.apple.com/research/introducing-third...
> unless you explicitly choose to help improve them.
Is it opt-in or opt-out? Big difference.
Explicitly choosing to help improve something by providing your data seems like an opt-in but what do I know these days
Opt-in = explicit and opt-out = implicit. You can't have implicit opt-in, that wouldn't make any sense.
You can't have implicit opt-out either. Opting is an explicit act.
The sentence prior states: "Privacy is the default, not something you have to manage."
So far Apple has been using default-checked checkbox during onboarding process for OS-level “help us improve…” diagnostics collection. Not sure about this specifically.
I don’t think that counts as an automatic opt-in, given that it’s presented to the user directly during the setup process.
It means my mother will have it on, at least, if they use the same pattern.