I don't know anything about this specific case but this exact thing will be a major barrier for autonomous vehicles to overcome. Instead of benchmarking AVs against the current, human-caused number of vehicular deaths - people will naturally benchmark AVs to 0 and treat every instance of death as a failure of AVs, even if the reality is that AVs end up 99% safer than human controlled vehicles.
As it should be. “Safer on average” is not a defense for failing something as basic as stopping at a stop sign. The bar for not running one and killing someone is not unreasonably high.
Safer than average is totally defensible, because every year the margin of what is "average" changes.
I'm just pulling numbers out of my ass here, but if say, one person is killed by an automobile every 100m miles, if the next year the average is 0.999 per 100m miles, that's a good thing provided the trend continues. We should not step over a dime to pick up a penny.
I do think there's a point where introducing different failure modes is harmful. Say they end up causing more pedestrian crashes and more head-on & angle crashes (such as the one in this article), but less loss-of-control and fixed-object crashes. That would make it safer for occupants but less safe for non-occupants of the self-driving vehicle. Since the occupant is the ones engaging self-driving mode they should be the one carrying the risk, so IMO even if it reduces number of deaths overall it should not be allowed
We do tend to impose them on individual drivers who make egregious and consequential failures, that’s the point. If every car running the same software is functionally the same ‘driver’ this class of failure absolutely should trigger a moratorium.
>We do tend to impose them on individual drivers who make egregious and consequential failures, that’s the point.
Do we? It's a common saying that if you want to kill someone with no penalties, have them run over with a car.
Then there are examples like these all over the place. Not to mention distracted driving and DUIs don't even get your license suspended in many or most jurisdictions.
>The Audi, which Jones had bought one month before the crash, was the third car he had totaled in a crash within 11 months. Speed was a factor in all three collisions, but police did not cite Jones in the first two crashes, prosecutors said.
Also see the 80-year-old woman in SF who killed an entire family at a bus stop, while speeding, while going the wrong way on a one way street, and was sentenced to probation and community service and banned from driving for three years.
No, because you can punish bad drivers, remove them from the driving pool, put them in jail, or in some way achieve justice (this may not always happen, but at least it's posible). There is no way for FSD to have any accountability, and so no justice can be served. "Sorry, your son was just unlucky that he was killed by the AI, but we can't do anything about it" is an intolerable breakdown of justice. Same problem we have with qualified immunity.
"Sorry, the driver could have turned on FSD, but decided to drive [drunk/tired/distracted] and killed your son. Thankfully we can ban them from driving for a while."
This could play out a million times at the population level. If you can decrease the chances of needing justice in the first place, shouldn't that be considered?
“Who to blame” turns out to be one of the most important facets of humanity, and it turns out we structure our entire society around it: assigning liability and pursuing justice. That’s not something that can be handwaved away to technocratically chase a metric (traffic deaths).
A person needs to die for us to assign liability and pursue justice.
If fewer people die, there is no liability to assign.
It sounds like you're arguing for a human being to die to preserve a human being to blame. Which yeah, sounds just like us! "This is the way we've always done it!"
If fewer people die, there is only less liability to assign. Remember that the original argument was "FSD needs to have 0 fatalities before we adopt it". And it's not valid to argue that some traffic fatalities are unavoidable; I know people who work at the DOT and "Vision Zero"[0] is a thing, and already implemented in some European countries/cities. Without FSD. So FSD needs to clear this bar.
> It sounds like you're arguing for a human being to die to preserve a human being to blame.
God I hate technocracy. There’s a reason we elect politicians and not technologists who think everything is simple and we just aren’t making the technically correct policy decisions.
“I’m sorry ma’am, we won’t be investigating anyone for your daughter’s death at the hands of the Sexy Deathmobile 5000 because technically no one is liable. But just think, there would have been so many other children killed this year if we didn’t allow these Sexy Deathmobiles to take over the roads! I hear there’s even a software update that will reduce toddler deaths by 5% at the expense of a 10% rise in dog and raccoon deaths, but it’s a choice our technocratic overlords are ecstatic to implement!”
You’re also arguing against a false alternative, which is the status quo. I’d rather have our money put towards safe streets and reducing demand for toddler-bulldozers (cars with raised grills) than a FSD pipe dream that would only enrich a few billionaires and the lucky SWEs who happen to pick the correct company to work for.
But you can take their license, fine them, jail them.
Corporations just pay a fine and move on. I don’t want that to be the standard for road accidents and frankly outside of this techie niche no one does.
It should be a critical part of the evaluation. If you set the bar higher than “safer on average” then you’re going to delay the switch, and every month of delay will cost more lives. “Safer on average” is exactly when you should start using the new thing, even if it’s uncomfortable and even if it’s a low bar.
No it shouldn’t. Low risk behaviors should equal low risk outcomes. If it’s safer on average yet causes low risk behaviors to suddenly become unpredictably dangerous, folks will very reasonably push back
Often we do not. The Freakonomics podcast had an episode titled The Perfect Crime where they basically say if you want to kill somebody, hit them with your car because that’s often not punished beyond a ticket.
>According to evidence presented from law enforcement at trial, on September 30, 2024, Burke was operating his patrol vehicle at a high rate of speed without emergency lights or siren activated when it crashed with another vehicle. Burke's partner was injured in the crash, and 34-year-old Cedric Hayden Jr. and 33-year-old Dejuan Pettis were killed.
Not sure why you are getting downvoted here though yes, the system is not perfect, in some states more so than in others. it is a fine line between giving someone who deserves it a 2nd chance vs. someone with 17 priors. but guessing you would also agree that allowing "self-driving" tech to kill people on the road without any repercussions isn't all that ideal either?
You can't compare autonomous systems that drive hundreds of thousands of vehicles to a single person. You have to compare it to hundreds of thousands of people.
American drivers kill ~100 people every day behind the wheel of a car. If autonomy were 50% safer and was deployed at such scale that it replaced 50% of all hours driven it would save 25 lives every day. It would also kill 25 people every day.
Sadly, many people will clutch their pearls at inevitable headlines like Bloodthirsty Elon Sacrifices 25 Innocents on the Altar of Profit Every Day! and, their emotions sufficiently primed, they will happily pull the plug on autonomy, sending 25 people per day to their deaths and feeling great about themselves as they do it.
Ok, in my country, you actually get your license removed and jail time if you kill someone (usually 2 years, up to 10 if you're drinking).
Let's say Tesla kill 50% less than a human driver, I'd agree to give Tesla's board 50% of the sentence for each person autopilot kills. As an incentive to make the cars kill less.
Let's reformulate it as a railroad problem. In the proposed hypothetical, if you effectively ban self driving cars until they're literally flawless, 50 people per day will be killed by human drivers every day. If you pull the lever and allow imperfect-but-better-than-human self-driving, 25 people per day will be killed by by autonomous vehicles.
As a reference point, non-suicide gun-related deaths total ~45 people per day.
What moral principle could possibly justify not pulling the lever?
But, even as everyone talk about morality as an ideal, no one really cares about it (else the world would be very, very different). What most people care about is feeling that things are 'fair', at least the appearance of fairness. I guarantee you that if a company accepts that it's board (or executives) get jailed if their car makes a mistake that kills someone, like any random driver, people will accept it.
Imho part of isuse here is that you cannot "partial AV" it.
You cannot have a car that will EVER need you to intervene or else it does not work because of awful human nature. The amount of tesla uber drivers ive had literally browsing tiktok, watching youtube UNTIL the car beeps that it is a green light detected, then they immediately hit the accelerator before looking up is ridiculous. You cannot have a half ass "autopilot" driving experience where you are expected to zone out and then be expected to take control within a millisecond during uncertainty. It does not help that FSD, AP, aAP, bAP, the model, whether the car is visual only all apply differently
Another huge problem is that error rate it not error shape. We have an enormous body of knowledge and process for guessing what might cause problems for human drivers, how to deter the problems, and how to mitigate the damage.
For example, you'll have a pretty good idea how our fellow humans will react to rumble-strips if I describe them or play a recording. I don't need to consult a datasheet or design an AI test suite with a zillion simulations and a bunch of closed-course tests.
What percentage of road accidents involve broken down old cars, drugs and alcohol, irresponsible and risky behavior?
Because most of us don’t have that problem, and insurance stats suggest we’re a hell of a lot safer drivers than a hypothetical aggregate including stolen car joyrides by teens and 50 year old teetotalers in a well maintained, modern car.
Safer than the highest level numbers isn't really a win though for autonomous driving. All human caused deaths includes drunk drivers, inexperienced drivers, and elderly drivers. An autonomous system should never be any of those things, so comparing them against is disingenuous at best and outright manipulative at worst.
That's extremely silly. drunk, inexperienced, and elderly drivers are all human drivers. You can't just exclude them from your baseline because they are not in perfect driving form.
Let's also exclude young drivers, stupid drivers, distracted drivers, drivers who never took driver's education classes, short drivers, barefoot drivers, glasses-wearing drivers who forgot their glasses, and one handed drivers.
The insane variation of human drivers is exactly why AVs should ultimately bring vehicular deaths down..which is the goal.
Use comments extremely sparingly. Most comments should be at the request of the user. When something warrants a comment, keep it to one or two lines: what the code does and why it's necessary. No background narrative, no replaying the investigation or failure mode, nothing a test name or the commit message already says. Applies to specs too. If a comment needs a paragraph, make the code clearer instead.
```
The comments Claude was leaving got absolutely out of control. Just lines and lines of LLM drivel that was barely intelligible and not remotely relevant to what a code comment should be used for.
Not sure I'm buying it tbh. I'm no fan of the American healthcare system, but we don't need to invent new accounting to make it look worse than it is.
Lots of businesses and industries have legal obligations to pay money for various things at various times, they don't treat that as pass through...it's revenue and expenses. Money is fungible.
Yes. Each industry has developed accounting standards that reflect the nature of their business. You couldn't run a bank or a payments company with a simple sales - COGS = gross profit model, it just wouldn't make sense.
In health insurance specifically, profitability is somewhat regulated and this gets at the accounting issue here. Insurance companies should maintain a medical loss ratio of 80-85% meaning that fraction of the premiums should be paid to providers. The remaining 15-20% is split between administrative costs and profit. Most of the article's forensic arguments around this are weak and circular and represent a misunderstanding of the accounting itself.
Should gas stations exclude the cost of the gas they're selling as revenue? There's probably a better argument to be made that they should be included than insurance companies be excluded. Unlike insurance companies, where the costs could come in randomly and over the span of months/years, the gasoline they're selling must be replaced (no randomness element) and is turned over in a matter of days. And if you think gasoline should be exempt because "it's not money", should precious metal or crypto traders get off the hook because those aren't money either?
No - this is a false equivalence. Transaction processing companies for the most part handle the in-and-out flows as a single transaction. Insurance companies hold on to the premium pool ("float") for long enough that they have time to realize gains from investing portions of it - the inflow and outflow are very separate.
That's not a good analogy. Stripe and Visa don't deposit the money in their account and hold on to it, it literally goes directly from the payer to the payee, they just facilitate the technical movement
Insurance money goes from the insured, into the insurance company's bank account, and IF the insured customers need services, it's then paid to service providers. If not, it sits in the insurance company's bank account as profit
Considering insurance premiums that are later paid as insurance claims as not being revenue is absolutely bonkers and there's a reason that's now how the accounting actually works
No, they're collecting money specifically on behalf of a 3rd party and then giving it directly to that 3rd party. They are custodians of that money only, it never even hits their bank account (goes into a dedicated trust account before distribution), and they cannot legally keep it. THAT is an actual passthrough.
This is addressed in the first paragraph of the pdf, with comparisons drawn to other industries and financial instruments where such income is not considered revenue. One can of course disagree whether it should be accounted this way, but the concept is not outlandish.
“This measure, while a standard accounting metric, obscures the strong financial performance of financial intermediaries such as health insurance companies, whose revenues are mostly pass-through payments between insured individuals and their health service providers. […]”
It seems to me that this document is almost entirely an argument for changing the accounting rules because of this distortion.
Seems to me that the argument is really "are my premiums a passthrough to medical providers" and I have a really hard time answering Yes to that.
If they are, then what do we call it when my medical expenses surpass my premiums? Negative passthrough? Contra passthrough?
What do we call it when I pay premiums for a year, never use a dime of it, and then cancel my insurance? I don't get that money back, nor does it get passed through to medical providers.
Do life insurance companies consider my premiums to be a passthrough to my eventual benefit payment or do they count them as revenue?
One of my favorite things about HN is the way that you can find someone who has produced definitive proof of the non-existence of the public education system in an offhand comment.
Funding didn’t make them smarter, but it did buy the hardware to scale it. Other guys had the exact same idea at the exact same time like RankDex and IBM's HITS algorithm. Google just got the cash first to actually build the server farms to run it.
They used Stanfords existing infrastructure … I don’t think the grant they were on paid for it directly. Pretty sure the cluster was a resource shared by several labs
It's unreal how bad the initial rollout was between HTTP/streaming and stdio, bearer auth and OAuth. Virtually every client/MCP server pair had a different portion of that matrix implemented.
The prose on the page is very unclear. My best interpretation is that they want to continue supporting stdio but that they don’t want it to be its own special protocol. The obvious way to do that would be to speak ordinary HTTP (version 1.1? 2?) over stdio and to use the MCP-over-HTTP protocol over the resulting HTTP transport.
This would be more complex to implement for a simple server, but it’s not exactly difficult.
I’m not really a fan. But if you’re building a protocol that needs to map to HTTP anyway, then maybe using the HTTP binding everywhere is not totally awful.
In the flip side: I’m currently designing an AI-adjacent protocol, and it will be able to map to WebTransport, but I don’t plan to define non-WebTransport HTTP bindings unless a very compelling reason appears. The main implementations will not use HTTP at all :)
> if you’re building a protocol that needs to map to HTTP anyway
I don't think there is any guarantee that HTTP will always be involved. For example I might be calling a local LLM via CLI/script on a server with a stdio MCP connector that just runs other CLI commands, and never sends any HTTP traffic.
Right. But there is a lot of real-world usage of MCP-over-HTTP-over-the-Internet, and a lot of “harnesses” want to support that use case, so they’re stuck either implementing the HTTP-based protocol or using a shim.
1. We already use it plenty, so we have lots of implementations,
2. it's good enough.
Cons:
a. what shall be the form of HTTP IPC URIs? hostnames for http: and https: scheme URIs are kinda out of place in IPC applications (we need something like sys.ipc.arpa, d-bus.arpa, etc),
b. the overhead of HTTP is annoying -- any decent RPC can be significantly more efficient, unless one uses HTTP/2, and maybe even then.
(a) makes me want to write and submit an I-D for HTTP over IPC by using such names as in the parenthetical above.
Those are mostly at a different layer. You can speak Thrift or Avro or Protobuf over stdio or HTTP or TCP or carrier pigeon.
gRPC spans layers, and it uses HTTP in a more intrusive way than even MCP does — it expects to own the entire URL space at the IP/port in question. Using gRPC in a nontrivial way for MCP would be fairly heavy-weight: you would probably need to set up reflection and figure out how to bind all the MCP calls to it unless you just use it as a tunnel.
But like, what is it? Odd that this reached #1 on HN. The README is pretty bare outside of installation instructions and a link to "Cordis", which is "A Meta-Framework of Spatiotemporal Composability." and "under active development. The API is not yet stable and may change without notice.".
New coding harness that seems to have some novel concepts and one of the pretty cool things on their landing page for it here: https://deepseek.com/harness/en/ is the Every Run is Traceable view:
"Everything the model sees is recorded in an append-only session log: system prompts, reasoning, tool calls and results, subagent scheduling, and every context injection. In the Trajectory view, you can inspect these records by source. Resume, fork, search, and replay all operate on the same event stream."
Seems pretty helpful - have sort of wanted something similar (I use Pi).
They also released this research paper that backs their whole plugin composability system that seems pretty cool: https://github.com/cordiverse/paper
Which are unavailable with the leading American models. You can't look at the complete traces of OpenAI or Anthropic model agents, as they are encrypted (there's been discussion of a couple of different ways to expose those, but that violates terms of service, and well, you shouldn't have to find complicated ways unencrypt your own usage logs).
I'm glad they're doing this also and that more people are adopting it. Event sourcing [0] is the right way to represent informaiton like tool calls, user interactions, etc. --- it makes it easy to fork conversations and maintain a cohesive conversation stream and stable message history that does not break the cache.
Quoting it in full so you don't stealth-edit your comment:
<quote>
I'm glad they're doing this also and that more people are adopting it. Event sourcing [0] is the right way to represent informaiton like tool calls, user interactions, etc. --- it makes it easy to fork conversations and maintain a cohesive conversation stream and stable message history that does not break the cache.
</quote>
why dont you mention that its your site instead of prenteding like something you discovered?
A harness is any wrapper around llm calls that manipulates llm interactions to achieve the process for which it is designed. Claude code et al are just one type of harness, focusing on writing code. The kind of UI used to interact with the harness and underlying models does not matter.
I thought the harness was mainly a TAI (tangible AGENT interface). Its a harness for the agent, not a user interface. That is bolted on top of the harness.
Fantastic tq I feel stupid because I’ve seen that for a while now. Do you know more about the autocompaction process? I could use some help there too with ghcp. I don’t see my context going up while abusing smaller context models with long inputs nor do I see a compaction indicator after turns.
You don’t think it’s because titanic battles are interesting and here’s a company that (a) gives you the weights to a frontier model for free, (b) publishes great papers with LLM architecture innovations, (c) is insanely cheap?
reply