Hacker Newsnew | past | comments | ask | show | jobs | submit | stavarotti's commentslogin

> I’ve been running Kimi K3 alongside Claude on my normal coding work, and for all practical purposes I can’t tell them apart.

If the author is here, I'm curious what this means. How are they running Kimi K3? Are they using pi, opencode, claude, codex, or kimi-cli? Is speed a concern?

Without knowing how the comparisons are being made, it's hard to agree that one can't notice the difference. I do.


I’m curious, shouldn’t Mythos have discovered this? At this point, based on all the marketing from Anthropic, I’d expect all software from them to be flawless given all the capabilities Mythos possesses.


PDFs are both awesome and terrible at the same time. I've seen screenshots of emails added to pdfs alongside tables that span multiple pages. Because you can do almost anything and guarantee that it'll look the same regardless of how or where it's viewed is a big selling point for a lot of businesses. It's this flexibility (i say madness) that makes PDFs notorious, and why some labs have document parsing as a leading product (see https://mistral.ai/news/ocr-4/).


I'll be using it tonight but grudgingly so. Grudgingly because after July 7th, I'm not going to all of a sudden, start paying API prices (and maybe that's the problem) when I'm used to a subscription that gives me multiples in comparative value. Perhaps this is the fabled "token economics will come for everyone this year" that I've been reading about? In any case, I'll use the hell out of it to extract as much as I can, then back to the trusted partners Opus 4.6 and Sonnet 4.6 (for however long they remain available).


Won't using it eat up the whole quota immediately forcing you to pay API prices anyway?


The token quota is completely unpredictable and changes month to month. Anthropic has a real penchant for riding the fine line of useful and dark patterns that make me want to write them off forever.


On the xhigh effort level (not ultracode!), and at the beginning of a new 5h session, I asked it to review a branch that has 300 lines changed/added, and went to grab a coffee. When I was back after a couple minutes, I saw that it decided to create a dynamic workflow with 60 something agents and hit session limits on my max plan.

When I'm subscribing to their plans, I have the expectation that I'll be able to get some reasonable amount of work done. These days, this expectation was already not being met for me with Opus, and Fable acting like this was the final nail in the coffin.

I cancelled my plan and I'm looking for alternatives.


For a period of time - then you go back to opus 4.8 or new sonnet 5.0 like some kind of AI pauper. Shine your shoes for some fable tokens g’vnor.


I am fully expecting the rollout of a Max 350 plan after July 7.


I locked my default model to opus 4.6 around the time of the nerfs. Such better results compared to 4.7+

That's enshittification for ya I guess


The claims of 4.6 or 4.7 being superior genuinely make me laugh. Adapt your workflow if needed and use the superior model instead of just kneejerk believing they actually enshittified a model with zero evidence except vibes on an undeterministic model output. Jesus.


4.6 was the last model that let you disable adaptive thinking and set max thinking token budget. I liked having that available, and still use it sometimes.


Your vibes are definitely better than his vibes.


What about all the benchmarks that show improvements in each generation?


Many of the improvements are the result of agentic loops and an emphasis on autonomy. Some of us don’t like that because the models go rogue and ignore design patterns, architecture, coding guidelines or other things that are important.

My friends and colleagues that like the agentic autonomy don’t care about the code, they feel like if it works it works and if an AI system is the only intelligence able to understand it that is ok.

I still want to be in the loop. They don’t.


the more agentic focused the better though?

sonnet 5 is very noticeably a much better model than any opus that ive touched

it actually does the things i want it to, and uses tools and triggers skills appropriately, vs trying to make stuff up


Agentic coding should absolutely care about all the things you listed.


It doesn’t for me.


Like another commenter said below, last Opus version to respect adaptive thinking and token budget flags was 4.6


It was quite clear 4.7 was a dumbed down high efficiency model they put out in a rush to handle the capacity issues they were having at the time. I've experienced myself substantial degradation on basic reasoning tasks, which were fixed in 4.8.


Bro, it's all vibes.

Models get dumber during the day and smarter during the night, I swear.

but I'm not willing to scientifically verify this, so I'm just going to go off of vibes- just like everyone seems to be doing with projects.


These vibes are pretty obvious even with casual use. Weekends are so much better.


In my case it that I'm tired and more likely to miss issues or mistakes. My idea of good enough is at a much lower level when it's 10pm and I'm about to knock off and go to bed in an hour.


Exactly.

The real step change I've seen lately is in the amount of complaining people are doing when their models aren't giving them what they want.


4.8 is much better than either of them as well.


I’ll continue to use the last great reasonably affordable duo from Anthropic: Opus 4.6 for planning and Sonnet 4.6 for implementation.


I just finished reading Incorruptible and a central theme (Anthropic is a case study) is that trust is singularly the most important currency a business has. The past few weeks have done wonders for Anthropic’s marketing but just as much if not more damage to the trust factor. Businesses will continue to use Anthropic because it’s the default and accessible where it matters (AWS, Azure, GCP, Databricks, Snowflake, etc). But the trust factor has dropped. It’ll be interesting to see if they can turn the tide. Maybe Fable will be too awesome for people to care about the past few weeks?


To be honest, given the overwhelming (and unfair, and unreasonable) pressure against them from the Leviathan, which the other companies do not have to deal with, they’re doing pretty damn good. In my mind the trust has actually increased that they can handle bad times and still push forward.


There is no reason to have less trust in Anthropic. It's not clear they did anything wrong. It's more likely the White House simply tied itself in knots, consistent with the last year and a half of chaos from them.


> It's not clear they did anything wrong.

Fable will literally sabotage you if it thinks you're trying to compete with Anthropic.


It’s the thousand cuts problem. Look at the stories over the past couple of weeks: silent downgrades that they then walked back, billing errors with claude code, highly sensitive classifiers that made it impossible to do simple things (I asked a few botany questions and my very long chat got lobotomized), and several more. The ban is only part of it. It’s the whole rollout and the fact that it won’t be available in subscriptions in a week or two.


It's a good reason to not tie your company's success to US based hosted AI though. I've started experimenting with GLM 5.2 and other than the tooling needing a lot more setup once you're there it works pretty well.

I'm hoping that some relatively cost-effective self-hosting solutions come about as a result of Hopper hardware being sold off as they're retired from DC use.


Perception is a lot more important than reason when it comes to trust. Whether or not we like that


Most people don't care about trust anymore, we live in a low trust society where this is to be expected. People gladly line up to be poisoned by fast food restaurants and trade 1/3 of their life for pieces of paper on a daily basis.


Silicon valley may be a low trust society but I havent given up hope on the rest of it yet.


After COVID at least in the West, I have.


Or at any rate, the "society" being gauged for trust-levels has to be something substantially finer-grained than a US state.


it’s still significantly far ahead of openai. gpt 5.6 looks like “better 5.5”. fable does not feel like better opus 4.8.


We have no idea what GPT 5.6 is going to be.

And 5.5 still ranks higher than Opus 4.8.


we’ve got the benchmark graphs from the announcement, and i hear tepid opinions from people with gpt 5.6 access. fable is far ahead of gpt 5.5 (my prev daily driver) also. much less passive aggressive / overly defensive.


I'm not subscribed to either right now, but have always preferred the GPT "personality." It does what I tell it to, not what it imagines. With Claude I was constantly jamming the escape key -- "STOP THAT".


>> The past few weeks have done wonders for Anthropic’s marketing but just as much if not more damage to the trust factor.

I don’t agree with this at all. IMO Anthropic has shown that that are willing to take even significant financial hits in order to stand up to their values and mitigate what they consider to be dangers and risks. Some people don’t like that or think it’s just marketing. But that’s exactly what Incorruptible is about: companies that are willing to take a stand, even in the face of overwhelming pressure from competitors, shareholders and naysayers.


This is assuming the whole "AI safety" thing was anything more than Silicon Valley kool aid. The government just bought into the marketing and radical safety woo woo wholesale and panicked.

You could legitimately argue this is a unique situation, a brief window where cybersecurity is being disrupted by new harnesses + a strong model. But that will be fleeting as other models and products adapt very quickly, and the long term benefits of keeping it from the market are questionable at best.

It's not a coincidence the export control was dropped after Dario (who is a hardcore AI safety activist much like Ilya Sutskever) was replaced by Tom Brown in the government negotiations.


These style of comparisons are decent at showing capability but they don't really show me what I truly want - a sounding board and implementer with senior engineer-level execution. When I look back at all the teams that I've been part of, the best outcomes came from white-boarding (sometimes in the metaphorical sense) with one or two people, at times arguing, then finally compromising on a plan. Instead of synthetic benchmarks that try to be objective, I wonder if there's a way test this, or maybe I'm opining on a way of working that will soon be gone?


> On novel work:

> Work that introduces new methods, highly creative ideas, or solutions that have not been used or experienced before. More generally, an approach that introduces an innovative strategy to solve a complex problem.

Something that I've been thinking about for the past year or so is coming to grips with the fact that the vast majority (anecdote) of software engineering work is not novel (and maybe that's okay). Few opportunities lend themselves to doing truly novel work. Other than infrastructure work and highly specialized software, pause and ask yourself when you last encountered software were you said "how the hell did they do that?" or "damn, that's nice" (for me, the most recent was Ghostty). I think much of the angst that people have when they fear for their job is coming to the realization that LLMs can do most of the "standard" work that a lot of highly compensated individuals currently do. We've built livelihoods around this and the threat of that coming to an end is genuinely frightening.


> I think much of the angst that people have when they fear for their job is coming to the realization that LLMs can do most of the "standard" work that a lot of highly compensated individuals currently do.

Amd do it better in most cases imo. Which is also hard to come to terms with, because there is a good bit of elitism/entitlement going around. The idea that a SWE is working at a higher level, which is beyond the reach of mere mortals, so therefore the high compensation is justified. Meanwhile everyone is, for the most part, doing some slight variation of the same thing as you suggested.

After starting out working minimum wage jobs I've always thought that the work gets easier and easier from there. Compensation and hard work are negativity correlated.


This is spot on ! Most of the work we really do is pure boilerplate and should be automated. While there are instances of interesting work those are far and few in between . The most recent instance of "how the hell did they do that?" for me was duckdb.


All public school students in the US at least are taught how to do basic scientific research. They should be making novel discoveries every day. The only thing that stopping them now is their own laziness.


> Something that I've been thinking about for the past year or so is coming to grips with the fact that the vast majority (anecdote) of software engineering work is not novel (and maybe that's okay)

Correction, essentially 0% of software is novel. Git wasn't novel. Chromium wasn't novel. Linux wasn't novel. Even C when it came out wasn't novel. Likewise Unix. They're all permutations of either prior knowledge, or evolutions of already existing concepts. They only might _appear_ novel to people who lack the depth to see what technology really is. Effectively applied physics (which has been solved for... over a few centuries at this point?) which itself is applied mathematics. There is novely to be found in physics and math themselves, but it's far out of scope of practical engineering.


> pause and ask yourself when you last encountered software were you said "how the hell did they do that?"

Like every month for the past 5 years? The progress in machine learning is dizzying. It is astonishing what can be done now with text, images, audio, video, code, etc...

If you don't study it, however, you have no idea how it works or how to do it yourself.

oblig. xkcd https://xkcd.com/1425/


The movie was great. I'm always fascinated by adaptations, especially what they choose to exclude. I thought the movie struck the right balance.


Such a fun game. It has the right amount of whimsy. Those shifty eyes…


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: