The 512GB Ultra is amazing, sure. But who is it for exactly? VC funded big spender founders? In that case why would they need local AI? The ultra rich enthusiast? But there can't be too many of those. So who actually buys these?
I think there's some mental inertia around what a computer is and what it's worth. This thing can build custom software for you, mostly autonomously. It can monitor things happening on the internet that are relevant to you, in a holistic and flexible way. We have one crawling the web for local events we'll like, and it judges them based on what it knows about us, and it tells us about the best matches every weekend, which has yielded some awesome outings we wouldn't have known about. It reads the literature on a subject in seconds and uses it as context to help in decision support. It's not the same value proposition as a computer 3 years ago, where most people are mentally anchored on what a computer should cost. Having it at home means that you can use it as a personal agent that always puts your interests first, regardless of what ad model the commercial providers decide to put in, and you can stash in it your medical data, what you buy, what you make, your worries, hopes, and dreams, without worrying about that being used as training data, or worse, something to exploit you commercially. I think it'll become considered totally reasonable to consider spending the cost of a small car on a computer, for many families.
Also, a lot of companies are looking at how to run capable models locally to cut some of their (massive) cloud AI bills. An easy answer is worth a lot to them.
By this kind of logic we'd paying something like ~$10,000/month for our internet connections. The perceived value needs to exceed the cost but that does not make it the only factor to consider price with.
What makes this expensive & sell well is it's not very fungible at the moment. Where else are you going to get 512 GB of high speed memory with a well supported accelerator attached that you can throw in the corner of anyone's home and not really have them notice? There are plenty of lesser options, plenty of noiser/power hungry options, plenty of harder to support options, but not really something in direct competition at the moment. Even the next rounds of the integrated AMD/Nvidia solutions are only targeting 196 GB of much slower memory and compute.
I think the difference besides the supply crunch there is that everyone connected gets the ~same internet, just faster or slower. Quantitative, not qualitative difference. On the other hand, a computer that can run Gemma 4 8B versus one that can run DeepSeek Flash are different enough experiences that I'd say they're effectively a difference in kind. It's been a bit since we had such serious stratification in outright capability in computing, rather than just how long it takes to get something done, or how many of something it can serve at once. In the early 90s, I think there were a lot more of those "this computer can do this thing, this one just can't" scenarios.
Closest competition I see right now are stacks of 2-4 connected DGX Sparks, similar lowish speed high mem, and about the same cost/gig.
If it were about capability instead of speed then you're welcome to pay me $5,000 for DeepSeek Flash running on an SSD :D. For $10,000 I'll even give you something which can run 700 GB models comfortably from RAM - not that the speed should be worth much.
Haha right, well, usable speed. And I guess there are some parallels with internet, the internet is technically usable with HughesNet, but people used to fiber would probably consider it unusable.
I like this train of thought. The inverse is saying that the cost of this computer is the value we give away to AI companies by doing compute on their servers with our data. And to take it another way, is the value to you, the cost of a small used car?
> And to take it another way, is the value to you, the cost of a small used car?
For me personally, not quite that valuable yet, but I think it's getting there quickly. Deepseek V4 Flash massively increased the value of local AI to me, to the point where it's displaced most of my Claude Code usage, its upcoming vision enabled version should bump it further, and it's only going to get better from there.
It's a lot faster, but a lot of it is also feeling free to discuss things I wouldn't be comfortable sending to Claude, with the idea that that info is now theirs in perpetuity. I got my genome fully sequenced recently (it's cheap now!), and I get a battery of blood tests every year. Wouldn't do processing on any of that with Claude, but local AI? Totally great.
And if I was running a company with a large cloud AI bill, I'd probably buy a wheelbarrow full of these macs. Cheaper, but also a more solid/predictable base to build on.
Hyperscalers later will move to their own chips. Nvidia would follow Apple strategy of selling the hardware. Nvidia would like to have near frontier open-weights model and sell the hardware, otherwise that market will be ceded to Apple hardware of M6 Ultra and future versions.
Sure, it’s pretty dumb/naive implementation, would need to be a lot more efficient to scale. Basically had Claude write polite/low touch dumb crawlers for a bunch of local sites (library, local events spaces, luma, theaters, maker spaces) and whip up a little frontend to let our family and friends manage a little text description of what their family members like, constraints, that sort of thing. Once a week, the crawlers look for new events, add to a db, and then run through the list and ask our local LLM to grade each event, given the text description and constraints. Take the top 20ish for the following two weekends and email out. It’s been super helpful, lowers the activation energy to go to more local events.
Probably don’t really need it, just makes it easier. I tried it with some potato class models first and they would ignore or mishandle constraints, especially when vaguely worded. Sometimes it would try to send us to events that were obviously in the middle of the school day, and I’d look at their reasoning traces and there’d be some really boneheaded mistakes in there. If it reliably sends you a decent fraction of garbage reccs, the emails stop getting opened, my SO has no patience for that.
Weird swipe, I'm not trying to sell it to you, this is just where I see it going, and why these things are going to sell (and why 512 gig M3 Ultra Mac Studios have skyrocketed in demand/price). It's not been an empty promise, in that it's been providing lots of very concrete value to us already.
Members of my C-suite are doing this today on a raspberry pi, so the promise is not empty. Key difference is they aren't using local AI to do it.
For the M5 Ultra, I suspect it would be valuable for someone who wants to achieve all the above and more, but with local AI due to data privacy concerns, and also not regulated data that comes with lots of other requirements solved by more traditional approaches.
Three possibilities:
1. The type of person who deals with lots of intellectual property using expensive Mac-only desktop applications that aren't meant for servers, whose mind has formed positive associations with the term "Apple Intelligence", whose values overlap with Apple's lawyer's values, who actually stands to profit from having a Mac that's more powerful than anyone else's Mac, whose long-term goals are not impacted by planned obsolecense on a piece of computer hardware costing over $25k ($50k after 1TB SSD add-on).
2. Trust fund beneficiary who wants to show off, LARP as #1, prime target for Apple's marketing.
3. 2026 kit for billionare-class iPad babies. All brain rot content is 100% local AI-generated. Never have to speak to your children again. A true "we have dead internet theory at home" machine.
If you actually want to learn, highly recommend a visit to https://old.reddit.com/r/LocalLLaMA/. People there are running stacks of 2-4x DGX Sparks to get to similar levels of memory, at similar cost.
If I can get Sol level capabilities on a $20k machine, then it is well worth it for my employer to buy me that machine for work as a workstation. When you start paying in tokens vs subscription costs due to enterprise agreements, you really start to see how much cash utilizing frontier models at the frontier costs (and I'm efficiently using luna and other models where possible!)
Even using multiple windows in parallel for as many as 5-10 hours per day, I find that I am not fully using my claude max (20x) and chatgpt pro (20x) accounts. I can for sure use up the claude max account, but chatgpt either gives me a free reset before I run out of tokens or I just fail to use the full quota. The quota for Sol seems like 10x that of Claude Opus at the same level, and forget Fable, you can use a 5 hour quota in 20 minutes.
But lets do the math:
Lets say a 20k workstation can run 1 inference at a time at the same speed you get with Sol hosted by openai (big assumption) and run an equally capable model (big assumption).
Each month this gives you about 100-170 inference hours on a Sol 20x Pro account, and 720 hours (if you utilize 24/7) on the workstation.
Assuming a 36 month amortization before the workstation has to be replaced due to no longer being able to run frontier models or is too inefficient due to electrical costs or what have you:
The monthly workstation cost is about $550 capex and $150 electricity -> $700/month
You would need about 6 Pro accounts to reach that capacity, which would cost you $1200 a month.
But this fails because:
- You most likely can't utilize the workstation 24/7. Your work hours will be concentrated into 6-10 hours per day.
- During work hours you are capable of utilizing more than 1 concurrent session. 6 Sol accounts would support as many as 20-30 during working hours, not all the time but if you could burst to that many (don't forget sub-agents and agent directed parallel agent workloads).
- In 1 year the cost of Sol level models is likely to cost a fraction of what it does now.
this leads to:
Workstation 1 Sol Pro 2 Sol Pro
Monthly cost $700 $200 $400
Raw capacity (hrs) 720 120 240
Usable capacity (hrs) 100-130 120 240
Concurrent sessions 1 3-5 6-10
$ per usable hour ~$6.00 $1.67 $1.67
Usable hours per $700 ~115 ~420 ~420
One of the advantages of LLM's is that you can set up a task list and tell it to burn through those tasks overnight. It's a rare night that Claude isn't busy for me, and I do burn through my 20x subscription, sufficiently that I downgrade from fable around about ... now in the week...
"You most likely can't utilize the workstation 24/7. Your work hours will be concentrated into 6-10 hours per day."
I have agents running 24/7 doing research, in fact I would argue this how they will be used for most programming tasks in the near future. For chatting, I agree local inference makes no sense. But for tasks that run continually, I'm not so sure. Personal computers took a while, local inference will too, but I think it will happen.
Sorry for the snark, but are you trying to cure cancer? What could possibly need 24/7 research in our domain, that doesn’t need your input every 30 minutes?
Normal boring CS scientific work. Just running running my experiments, reproducing other papers, etc. A lot of it does involve the agent waiting for some computation, but the fact that it resumes independently when I'm sleeping is kind of the point (+ usually I have several running in parallel).
I'm not trying to cure cancer, although I do hope people who are use LLMs. ;)
> - You most likely can't utilize the workstation 24/7. Your work hours will be concentrated into 6-10 hours per day.
isn't the whole point of all this ..... agents? isn't that what literally everyone is always clammering about in these threads? in which case the workstation is useful 720 hours out of 720 hours.
Subscriptions are, and will likely remain, the best deal in town. Unfortunately, larger companies aren't able to do that. When your monthly token costs are in the $5-10k range, the local inference starts to look a lot more attractive
In the case where you pay for tokens without a subscription, the analysis is still very much not in favor of buying hardware.
The assumption previously used was that you can run a Sol level model on an M6 or whatever hardware $20k gives you. That is not true, it was an assumption made to show that even giving your own hardware every reasonable advantage it still loses.
Lets compare buying tokens of the best model you might run on your own hardware (still being unrealistic in favor of your own hardware) vs that same class of model on the market. I think one of the best models you might be able to run is GLM 5.4, but lets just look at chinese models generally:
$20k workstation, best case: $15k M5 Ultra 512GB, 36-month amortization, ~$440/mo. Runs a GLM-5.3-class model at ~30 tok/s. Saturated 24/7 it produces roughly 58M output tokens/month.
Buying those tokens:
DeepSeek V4 Pro @ $0.87/M $50
Kimi K2.6 @ $4.00/M $232
GLM-5.3 @ $4.40/M $255
Kimi K3 @ $15.00/M $870 (does not fit on the box)
The economics can never work in your favor for buying your own hardware here, unless you can utilize it or sell excess capacity and you have access to nearly free electricity. The reason is someone else can buy the same hardware at scale (or realistically more efficient hardware), park it somewhere with very cheap electricity, and sell tokens. They can get very high utilization that you are not likely to get.
And keep in mind I am giving 'your own hardware' no overhead or maintenance cost, despite your condition that it's in a large corporate environment. In reality corporate IT would make it almost impossible to set up and your would need huge lead times to buy the hardware and get it installed.
> $20k workstation, best case: $15k M5 Ultra 512GB, 36-month amortization, ~$440/mo. Runs a GLM-5.3-class model at ~30 tok/s. Saturated 24/7 it produces roughly 58M output tokens/month.
For agentic coding, ~90% of the cost comes from cached input tokens. This cost increases quadratically with the session length. If sessions go near 1M context, the number of cached input tokens can easily exceed 1B in a day.
So yes, at that speed for sure. But if the speed goes up? or the ability to batch at the same speed goes up? The economics start to shift. The gap is much closer, and you'd end up with a box you can still use or sell later.
As speed goes up the cost / Mtoken will necessarily go down at roughly the same ratio so it will wash out. The still use hardware or sell hardware value is factored in to the amortized monthly cost, it assumes a 3 year markdown, and does not factor in the cost of money which should almost cancel the resale value in the end, which I think is quite accurate (any residual cost on a graphics card after 3 years is so small compared to the current price it should be discounted and in included there).
Where you might win by owning your own hardware:
- Hardware costs go up, and thus api costs go up. You've locked in your pricing.
- Chinese/Open models become illegal/hard to access the way we do now. OpenAI and Anthropic are trying very hard to build a regulatory capture scheme to do this. I think they will be unsuccessful because China just won't participate.
Yeah, I'm not sure people realize how expensive ZDR/Zero Data Retention is, and how important it is to a lot of businesses, this kind of thing starts looking really cheap really fast if it's a reasonable substitute.
I think "ultra rich enthusiast" is in the right ballpark. There are people betting on being able to create their own revenue generating products and services with their own local hardware and very little operating costs. That may or may not make sense as a business idea. But people with wealth and risk appetite trying a new kind of business model and cost structure has a strong tradition.
Put another way: If $25k is the full extent of the start up capital costs, and operating costs are very low, that is a much cheaper business to start than most! The question is whether this is actually a useful model for a revenue generating business. I think that remains to be seen.
My guess is that there will be a few hits (which we'll hear a lot about - especially when someone actually pulls off "the first single-person unicorn", which I do suspect will happen someday) and a huuuge number of misses, which we won't hear much about.
Enthusiasts buying these for fun are not the target market. These aren’t big sellers to begin with but a lot of the sales are going to companies where people have budgets for gear like this and can make a business case for it.
This is, sadly, probably a foreign concept to a lot of people who have only worked at companies where hardware purchases are viewed as something to minimize and everyone is stuck with the same low spec laptops that the finance department picked out. At companies where someone might have a legitimate use for a $20K machine, their fully loaded costs (not their salary) are $300K or more, and other teams like sales are spending thousands of dollars per week on things like travel and hotels for their job, spending $20K on a computer that’s going to last several years is not a hard choice.
Other than the local AI crowd which is much recent it is professionals using Final Cut Pro for video editing, Logic Pro as a DAW and music production, Video transcoding, Photoshop and other tasks for high performance computing that don't need Laptops but want above 128GB of ram and prefer a Mac. Then there is the obvious group of developers that are making Apps for all of their products. Also, these are great for the workplace. AI is much more recent thing that Apple products were used for.
If Apple didn't sold these things they wouldn't make them but, also the level of marketing that Apple is talking about for AI is basically the new group they need to capture because the ones I just listed are already buying Macs and or easily to motivate with the other obvious CPU / GPU performance upgrades for code compilation, faster memory and video transcoding.
Local AI is the future, and a lot of people want the first mover advantage or to toy around with it. I know a guy with a small rack of Nvidia Spark machines that he uses for that purpose; it's as much as a decent used car.
But is it? If it’s cheap enough latency doesn’t matter. Unlike say cloud gaming, where latency does matter. I’ll take my games local and my text bots cloud
Of course. People are valuing privacy now, and it's almost inevitable that humans will seek to have more agency over a social contact, even if it's an artificial one.
The first three are starting to get more closely associated with privacy.
Convenience: if I want information from a chatbot, I don't want to have to hear how I can save on my car insurance by switching to another insurance company. That's basically what websites have become. Go to any American local news station's site. It's just wallpapered with ads and you'll inevitably get a popup asking to subscribe. I don't want that from my chatbot.
Cheapness: this doesn't matter over the long run because inevitably, both straight-up monetary payment and revenue generated by invading privacy become part of the service providers' revenue streams, if they don't start out that way. Cable TV used to be ad-free, as did streaming services. Then there was a need to fund the coke habits of some finance guys in Lower Manhattan, so ads were introduced as a "free" tier. Now there's only "reduced" ads on the paid tier, and you still fork over your data to let service providers give advertising clients a better profile of you.
Safety: identity thieves, stalkers, and even government agents acting against the law use commercial data sources to do things they otherwise couldn't. Imagine what they could glean from chatbot or other AI sources.
If you're your own LLM service provider, none of this is a problem.
I realize it's not a 512, but for context on why I ordered a 96GB. $95/month after a trade in. I'll use it for a long term project where I'm building a national sized dataset/processing video transcripts with a rubric (of churches/sermon health). I'm willing to trade SOTA models for local ones for cost of regenerating/model consistency/ethical reasons.
I'll find other ways to use the power though, opencode or maybe start working on more video projects.
My friend plans to get the 512GiB M5 Ultra ASAP. He makes money by selling synthetic data and training models for companies in San Francisco and it's the cheapest way to be able to do that on your own local hardware. It's also a bit irrational as he could save money by just renting, but I get the appeal of owning something outright.
The competition is a custom multi-GPU NVIDIA RTX pro desktop, which go for much much more. $20k is cheap for 512GB addressable memory. The old mac pro could easily be configured to cost that much.
Aside from high-spending users, I think many people use it as a productivity tool. If you can use it to make money, and the money you earn far exceeds the monthly payment, why not make your work smoother?
Anyone dealing with sensitive or regulated data can get started instantly with local models where they may need time consuming process or complexities to send sensitive data outside.
this comment has always existed behind every apple release, most especially anything vaguely pro-ish.
to answer your question : looking at the aftermarket availability of Apple's prior best and brightest : practically no one buys them.
"people here buy them" , well, 'here' is one of the most affluent groups of people in the world.
They're available as movie and television set pieces (undoubtedly disappearing into the home of someone close to the staff post-production), and for administrative/boss types that can slip the cost into a ledger somewhere that few will ever see.
It has been a hobby of mine every few years to check out the apple site and see how big I can option a machine. My record was when I was in high school years ago and was able to option some pro studio-ish apple desktop thing to like 61,000 usd out the door.
Movie set pieces, as a motivation for Apple making these high-end configs available? That makes no sense.
For one thing, you can’t tell from a movie what the specs are. A $999 Mac Studio looks exactly the same as a $20,000 one.
For another, Apple updates the industrial design on their products so rarely, a 6-year-old Mac, iMac or MacBook also looks nearly indistinguishable from a brand-new one.
I don't mind no windows at all if they have "digital windows" as substitutes. You're already 40,000ft in the air I think windows aren't really gonna help your claustrophobia. Not like you can open the door and walk out for a bit.
Maybe. I think people could survive it, sure. But I also wonder if there's a part of our lizard brain that just for safety's sake prefers to be able to see out of any small enclosed space we might find ourselves in. And that lizard brain is smart enough to know that a flat screen is fake. Again, I just think of flying in a plane today but with no windows and I, at least, am immediately squigged out.
I also feel like a lot of modern air travel is designed around making something that's highly unnatural and inherently terrifying into something palatable for most people. I'm not talking about actual safety, but the perception of safety. Which isn't exactly a logic thing. More about keeping that lizard brain from freaking out too hard.
I get motion sickness when sitting sideways or backwards on a car. People tend to become nauseous when they sense motion but don't see it, or what they see doesn't align with their other senses. Even the small amount of lag that a "digital window" would likely have, will make things much worse.
Do you think the barbarians are at the gates of OpenAI and Anthropic? If cheaper, open weights models can seriously take revenue away from those two labs for (frontier - 1) model use cases (which are the models most enterprises will choose) then OpenAI and Anthropic are left only with users using their latest and greatest model AND who will keep upgrading to the newer ones?
The "singularity" as stated requires AI to either make a technological advancement strong enough to be deadly to humans (besides just intelligence), or spreading deep institutional support for itself among society.
The idea that "we will get superintelligence first, then... ???" is kind of a weird notion. I mean, it's pretty arguable that we do have at least some form of superintelligence. The AI itself needs to actually do something with it though. Either that, or more likely, someone needs to do something bad with the superintelligence.
That could be both re-assuring or not. Because under that view, given how AI is being integrated so quickly into society, it's not going to take this fantasy view of superintelligence to reach the singularity. If you have broad institutional support and crowd out the thing we call humanity over time (the two ways to 'solve' a problem: solve it, or declare it meaningless), that is another way to reach the singularity.
Without the singularity, I think Frontier labs will offer intelligent model blends. They'll have their own versions of "cheap" models and be expert at using the appropriate amount of compute for a task.
Is that impressive? If an LLM is just spewing out code which it has been trained on extensively then to me it's not that impressive at all. I want to see individual developers creating new software which previously took a whole team. Things not done previously. I am very pro AI btw, but the constant barrage of false hype is really tiresome.
One place where fiber cables cannot reach would be... way up in the air. Think about how many people fly each day and then remember how poor internet connectivity and speeds are at 40,000 ft.