I know. It hurts. But I have the feeling that you get what you pay for.
The purchase cost of H100 or B200 systems with comparable VRAM is a one order of magnitude higher. Although I can only guess how much lower the token/sec output of the Mac Studio will be. Probably 2-3 magnitudes lower?
While a cluster has to work with many users simultanously, and is a good investment for a company, perhaps the Mac Studio will be a good use case for a personal larger LLM deployment configuration.
Perhaps someone has the token/sec numbers for larger models running on older Mac Studios?
Back of the envelope is that compute doesn't matter for inference, only memory access speed matters. Tokens/sec is going to be in the ballpark of how much time it takes for the compute element to read the entire model. So you can give GBtok/sec (sec = seconds, GB gigabyte size for the chosen quantization)
Q8 is a nice quant for this calculation, since 1 byte = 1 weight, so Qwen 27B Q8 will be the tok/secGB value divided by 27, and Deepseek V4 Flash 162B, by 162.
NVIDIA B200: 8,000 GBtok/sec
NVIDIA H100: 3,350 GBtok/sec
NVIDIA A100: 2,039 GBtok/sec
NVIDIA RTX 5090: 1,792 GBtok/sec (deepseek only fits at Q2 or less)
Apple M5 Ultra: 1,200 GBtok/sec
NVIDIA RTX 5060 Ti: 448 GBtok/sec (deepseek doesn't even fit at Q1)
Apple M5 Pro: 307 GBtok/sec (can do deepseek only at Q4 or less, at maxed out sped)
Apple M1 Pro: 200 GBtok/sec (can do deepseek only at Q2 or less, at maxed out sped)
Apple M6: 170 GBtok/sec (can do deepseek only at Q2 or less, at maxed out sped)
So a base model apple M6, Qwen3.8 Q8 = 170/26 tok/s = 6 tok/s (or 12 tok/s at Q4), and a B200 will do ~40 times that, or 240 tok/s. Which is kind of sad as a base model M1 pro will beat it comfortably despite 6 years of chip advancements. To add insult to injury an M1 pro ... is cheaper secondhand.
Liquidity providers like Jane Street, Citadel, et al make money on the spread. They also buy order flows from integrators, and retail investor order flows are now a product.
i.e. retail investor → brokerage platform → clearing/execution infrastructure → Jane Street → payment back toward the brokerage side of the chain.
Who captures the economic value created by retail order flow?
Jane Street.
In an ideal market, this product line shouldn't exist. Institutional investors should not be making money on the activity of retail investors.
What incentives determine where that flow is sent, and would investors receive better execution if their orders were exposed to genuinely competitive price formation rather than privately internalised by a concentrated group of wholesalers?
The regulators should be squashing any HFT related or retail order flow, but it's so opaque _by design_ that getting policymakers, or the general public, to understand that retail investors are paying some portion of tax on their $20T USD annual trades to these companies.
Granted, these order flows _sometimes_ work the other way -- and retail users get a better deal on a trade.. But would you really expect the market to be worth what it is, if that was the case less more often than not?
There is a clear and obvious conflict: the broker is supposed to seek the best execution for the customer while potentially being paid by the firm receiving that customer’s order. How can that be, when the broker's in bed with the liquidity providers?
PFOF and HFT are distinct concepts, but they are widely conflated in this thread. I don't agree that PFOF is inherently bad, but even if it were: it is not a valid criticism of HFT.
Thanks for bringing this up and trying it out! We've disabled Cmd+Option+I in the new version and also have a message during signup about sharing chats with Bullet, let us know if you run into any other issues!
OP was providing a userful tips to users, and they weren't giving you a bug report to remove it. If anything, they were suggesting you remove a login-wall.
But now that you blocked the inspector tool, you'll find it harder for users to report bugs. Unless of course those are automatically shared too? /s
We're actively working on implementing a guest mode feature right now for users that don't want to make an account. For reporting bugs, we have an in-app feedback form. Bugs are not automatically shared.
I guess the question is; is this the late-cycle cash-out (aka harvest pricing) akin to Sun Microsystems at the dot-com peak -- or is it a repeat of the crypto pricing hijinks we've already seen from Nvidia, where they're just exploiting the lack of supply?
That's a 90% uplift since the original pricing, in a market that is already showing signs of seizing.
With Apple offering leasing options, CXMT knee-capping Samsung, all of the big tech players on a run to outspend on CapEx by the end of the year..
We're at a point where local models are exceptionally capable, model-on-silicon dies like Taalas (recently acquired by AMD) may be cutting inference cost substantially for the 90% of work we do day-to-day (similarly Alibaba's T-Head division with open-model-forward inference chips being produced domestically in China).
If we shifted all of the design/planning to cloud models like Fable 5/Sol 5.6 Ultra, and day-to-day operational inference to these chips -- it's quite likely we'll squash usage to single-digit percentages of what we're currently using. But we should also expect the model providers to take a similar approach.
In any case -- I'm keen on demand destruction, both as a consumer and someone without skin in the game.
It makes you wonder to what degree they are making more money by raising the prices and restricting supply absent meaningful competition, or, alternatively, by making them in volume until they can supply the total demand. The lack of competition is probably a big part of the problem, but then again, that competition might just increase the price as well if it existed. AMD really could make a killing if they got off their assess and properly supported their cards (and made them competitive performance wise, which is probably a really hard thing for them to do).
i think maybe by pricing up they want to make a reliance for people on their services rather than developing running their own. if these devices are out of reach for many it means less innovation and then less competition. nvidia is now tightly coupled financially to business who offer services that owners of such devices might replicate without using services of their partners.
In my experience, program requirements are mostly there for the lawyers.
If you act in good faith, communicate clearly, and conduct yourself reasonably, most companies will work with you — even when you've technically wandered outside the neat boundaries of their risk-appropriate, regulatory-reviewed policy.
That isn't protection, of course. Eventually you'll encounter a bounty program run primarily by lawyers, procurement, or someone optimising a graph trend-line.
And we know what tends to happen next.
Those programs, and organisations, develop reputations. Researchers talk. Companies get discussed at conferences, in private groups and across the community, and some become informally blacklisted.
Microsoft is a useful recent example: researchers have publicly walked away from five-figure bounties to make a point. There are excellent people working there, but organisationally Microsoft has repeatedly struggled to engage with the security community in a way that feels collaborative, rather than adversarial. Unless you're one of their paid partners, intermingled in their ecosystem.
A lot of that seems to come down to incentives: somebody, somewhere, wants the numbers to look better.
That doesn't work particularly well in an industry that, like most industries, ultimately runs on relationships, trust and specialisation.
If a company marks something critical as informational, sometimes the most effective response is a CVSS parameter argument. It's a snarky comment:
"Okay — so if I find a way to abuse your own infrastructure to message your customers, trigger a major incident and create regulatory problems for your clients, you'd prefer I treat that as informational too, and instead just report it to regulators?"
Surprisingly often, that gets the issue reconsidered.
Have you annoyed an analyst? Maybe. Does it matter? Probably not. Neither of you will remember the exchange a week later, but you might have corrected a bad risk decision on their side and you'll see a positive outcome on your side. Mistakes happen.
I generally advise companies and hackers alike to follow Kiwicon's #1 rule.
I've disclosed vulns across just about every industry — banking, healthcare, oil & gas, government, cybersecurity, etc -- and to some of the largest companies in the world, OpenAI, Salesforce and Google. I've been doing this for nearly 20 years.
Most of my research starts with: _There is absolutely no way this works_. Then it works.
I've been thinking that a lot more lately.
Companies and hackers are both heavily incentivised to reduce the friction involved in vulnerability disclosure, particularly for large organisations. The platforms are good enough now. They're email in 2007: imperfect, occasionally frustrating, but substantially better than what came before.
They make SLAs possible. They provide structure and administration. Things still go wrong — companies stop responding, analysts drop the ball, hackers can be idiots — but the model basically works.
Decentralising disclosure again would make life significantly harder for individual hackers. We'd end up back on email, probably building email-powered bounty CRMs that consume a small country's worth of tokens just to keep track of everything.
For smaller organisations, though, I wouldn't touch a public bounty platform with a 10-foot pole. Run a private program first (through the platform). Having been on the receiving end of beg bounties, automated scanner output and increasingly AI-generated slop, most smaller security teams simply cannot scale to absorb the noise.
The more interesting way to think about these platforms is that they're becoming the LinkedIn of hacking.
For hackers, the path is fairly straightforward: build a rep through useful -- but oftentimes unsolicited disclosures, get invited onto private programs, and gradually establish a profile with a strong signal-to-noise ratio.
For companies, they're increasingly a recruiting and relationship-building tool.
And for the platforms, I think there's a much larger opportunity for them in community.
They should be significantly better at understanding hackers: what they're good at, what technologies interest them, which industries they understand, and where they're located. Today, that profiling is laughably poor, to the point the questionnaires on areas by these large platforms are out of date by several years.
Then use the data.
Run small, highly targeted events: state- or city-based meetups, lunch-and-learns, product launches, bounty program launches and technical briefings. They don't need huge sponsorship budgets or prize pools. They need the actual community involved. Pay for dinner, sponsor a talk.
A lot of existing events seem to start with companies, sponsorship packages and monetary amounts, then work backwards. I think that's backwards.
As a weekend hacker, I'm far more likely to spend time on a program because something about it is interesting: you're launching an AI feature, handling financial data in a new way, using Node/GCP/a TI-82 calculator, or exposing some weird technical surface I want to understand.
And I'm far more likely to build a useful relationship with a company if I can actually meet the people behind the program. Hackers can provide much better feedback than a semi-generated report, and companies can explain far more than a stale domain list and scope document — which, realistically, we'll be ignoring 99.99% of the time anyway.. Unless it's government. I quite like my freedom.
Is this true? If so, how do you know? I have listened to almost of their podcasts. I don't recall them saying there are any type of customer they refuse to sell to. They told a funny story about a sales call with a US national laboratory. They went into the call assuming they would be asking for supercomputer. Instead, they learned they need a bunch of regular rack compute, not all supercomputers.
Also, the OP did not say they are a SaaS company. They only said they spend 900K USD per year with AWS.
I also think they work with government customers. I saw an open job position on their website requiring TS/SCI security clearance and full scope polygraph.
Data theft (i.e. taking corporate IP) was normalised until very recently (I'd say the last 2-3 years with the rise of the adoption of DLP and public litigation). A few examples I can think of thing I've heard _just_ in my own career; terabytes of client data taken by a former consultant, sales reps that joined _solely for the purpose_ of taking this quarter's leads and then ghosting the company, execs forcing the use of their buddy's startup/consultant/vendor (and then they later join the startup/consultancy/vendor). The latter being _the most common_. If you're a large company, your sales pipeline is so valuable to startups -- and from a single consultant perspective the ROI of ripping off your sales data is _insane_.
I think back to the questionable things _I've_ done that _don't_ amount to data theft, but _should have_ caused someone to probably kick off an incident/investigation.. But they didn't.
- git cloning every repository in the company (who am I kidding, I do this in _every_ company
- airdropping stuff to my personal phone (my profile photo, but it could have been the aforementioned git repos)
- using sharedrop/similar services (to copy my RSS/news feeds, favourites bars) from my work laptop
- enumerated staff/ops dashboards/tooling to get shit done (think: enumerating the CRM before Lazarus and Lap$u$ made it cool)
It used to be a thing that dev were keeping source code and all in old time.
I remember a manager during an internship telling me to do clean code because later I might have to reuse piece of code in further work experiences.
It was the time when code was shared with zip and so or net shares.
At that time there wasn't that many "software shop" and it was just a way to do things in other company like hardware manufacturers. Obviously you would have been trusted not to take or disseminate or reuse company trade secret or coffee things like expected by your non competition clause in your contact anyway.
Since then, a lot of company became like "software production" company and developers are now considered like factory workers.
And in some way I would say that you are nowadays robbed of your code and there isn't even attribution anymore.
Look, in the 80/90s it was more common to know the name of the main dev of major companies, and their contributions were clearly attributed to them. But now you will very very rarely know the dev that did the code for anything. Top manager/architect/... Might be recognized but not really for their code contributions.
You’re so right, I totally forgot how normalised it was to take common functions and classes and factories. I guess ultimately, it was all worthless anyway thanks to LLM’s.
I actually recall the lead engineers in my department coming into work on their last day with a harddrive. The same people who had access to national databases, and production systems that impacted the majority of the country’s population. That simply wouldn’t happen today.
Although I’m not someone who’s dealing with PCB schematics or bleeding edge IP.
I guess I’m a dummy, I always deleted all cloned repos from companies I ever worked with. I’ve always been way too terrified to keep anything company related
Oh this was on company issued devices, I don’t keep the source code! I was just explaining how this is something that probably should have triggered an investigation, but didn’t.
8800 GTX in 2006. Cutting-edge, an insanely powered consumer card for the time. Theoretically around 0.3456 TFLOPS.
1080 GTX in 2016. Cutting-edge, an insanely powerful consumer card for the time. Theoretically around 8.87 to 8.9 TFLOPS.
5090 RTX in 2026. Cutting-edge, an insanely powerful consumer card for today.
Theoretically around 104.8 TFLOPS.
In the same timeframe mobile processor CPU's went from 0.001 TFLOPS, to today's Apple's A19 Pro chip which delivers 2.074 TFLOPS.
That's _without_ getting into ASIC's, or purpose-built hardware like Taalas's model on silicon HC1, or generic AI dies like what they're planning with HC2 or Cerebras, which will massively compress the timeline.
Yeah but here you describing the opposite phenomenon. You're saying that the hardware is going to become cheaper and more powerful with the years, to the point a current State of the Art model from today will run on a normal consumer hardware in ten years. What people are trying to do now is the opposite, optimize the software as much as possible so that it does not need the best hardware but the normal one we currently have. As if we were trying to make a current AAA game to run smoothly on the 1080 GTX of your example.
(2) Are you saying that you think we're at the limits of computing in general, or that specific technology?
We know, for example, that a human brain level intelligence is possible to run on a human brain. We are nowhere near that. And actually that's not even a physical limit necessarily.
Leaving aside the discussion on LLMs intelligence vs human intelligence, on a purely energy consumption level we are definitely and without any possible questioning nowhere near that indeed.
Here you have shown yourself that progress slows down and doesnt speed up. 8.9/0.35 = ~25x more performance in 10 years from 2006 to 2016. 104.8/8.9 = ~12x more performance in 10 years from 2016 to 2026. Growth has dropped 50%.
That isn't deceleration, you've just chosen a very selective way to compare. If you use time as a denominator, which is kind of intrinsic when talking about rates of acceleration, you get a very different result. If you graphed .3, 8.9, and 104.4 on the y axis, with years on the x axis, it would be pretty clear that there was in increase in the rate of progress.
We went from adding 8 teraflops in a decade, to adding almost 100 the next decade. If we add "only" 400 more teraflops in the next decade the graph will make that initial growth look flat in comparison, even though your math would show that we are basically stalled out.
It’s like claiming that a company that goes from making $1 to $1k to $100k to $1mm in a 4 year period has decelerating growth.
We will see such power and price now only when AI market crashes or China reaches node parity and goes after market share as currently the way they are buying out most of the latest node production the consumer prices will only be palatable to the very rich or we will need to be happy with older slower nodes
Speaking of ASICs - how likely is it that as models get better we'll see someone baking a whole model directly into the silicon? It's like having l0 cache.
This is definitely being done with private models by HFT/quant firms, data processing agencies/orgs (large intelligence agencies, _every_ data analytics org, etc).
Yes, right now it would be obsolete in six months, but I also must add that this never stopped crypto miners from making new ASICs.
However, with how useful Kimi is right now - at some point if someone makes a dedicated hardware board with "good enough" model for daily tasks - that would be a very sought after commodity.
They're important everywhere of course, but especially on mobile. If AI researchers figure out how to offload knowledge and expertise from reasoning weights, then a core reasoning ASIC linked to the knowledge would totally rock.
Just a note that I think the direction most people are paying attention to is memory bandwidth; thats the real bottleneck and “number go up” but also constraint people are designing around