Hacker Newsnew | past | comments | ask | show | jobs | submit | aynyc's commentslogin

It means a language for the very elite special force unit. In MSFT, that means the Ads teams will be able to serve you Ads in excel very quickly. /s

Please execute my ignorance. And this isn't an attack on anyone, just my own stupid limited understanding.

Aren't JellyFin & Plex essentially just a catalog for your local data (most likely torrented movies, etc.) in a web app? If so, what's the big deal?

edit: ah, i missed the BIG transcoding part.


Sure its a catalog of your local media served via an API, with metadata enrichment, built in transcoding (hardware support too) and multiple user support (watch history, recommendations etc).

Its not just a web app, I would say most dont use the web app directly, instead they provide clients (Android/iOS/Roku/AndroidTV/WebOS/Tizen) that are in official app stores. This is the big thing for me, being able to just install the client on your TV and not worry about it is great.

Often the clients are just wrappers of the web client with native integrations, but they work well.

Plus there are loads of 3rd party plugins to extend it, and alternate clients, Im a fan of https://github.com/jeffvli/feishin for music duties personally.


They do a bit more than this. They do streaming and transcoding for a start, that's fairly essential for a functioning video platform. They also have native apps for a range of platforms. Support subtitles, multiple audio tracks, all sorts of codecs. And then there's all the metadata stuff as well, getting blurbs, posters, ratings, categorisations, episode naming (which has so many awkward edge cases) etc all from various metadata services. Oh and they do this all with multiple users and ACLs.

Plex also then adds social features and curation. The value of that to its users is debatable, but it's additional complexity regardless.

These are pretty big projects. A founder moving on is a fairly big deal.


I can browse to jellyfin on my phone, ipad, laptop, TV etc watch any format (since it will transcode on the fly) and watch whatever I was watching from wherever I left it. It will also obviously know where I'm up to in any series so if I start watching something else I can go back to it at a later date.


A friend or relative can stream the content over the internet via Plex or Jellyfin (through a Roku, or other streaming device) via a UI that is very, very easy to use. Plex is a way for my non-technical aunt to easily interact with my media server from her home three states away.


This is about their bond, not that share price. If you are in the US, it's like having low credit score, everything you want to do financially such as leasing or financing a car, buying a house, etc.. will be more costly (higher interest rate) from the lender.


Except that it's not clear really that lenders will be able to judge on that basis alone. SpaceX is a different kind of unicorn: it's a government contractor run by the richest man in the world who controls a media echo chamber and gets people elected.

That article compares them to Oracle. Who are, as it goes, pretty similar: run by rich people with a media empire who have their teeth deep in government systems.

These bonds could get worse and worse but if US state and federal governments continue to put thumbs on scales it doesn't matter. The US free market isn't uniformly free.


He's only the "richest" man in the world because of the inflated valuations of his companies. At some point the market looks like a dog chasing its own tail:

Q: Why is this stock so valuable? A: Duh, because Musk is worth a kajillion dollars and everything he touches is gold!

Q: Why is Musk worth a kajillion dollars? A: Because he holds so many shares of extremely valuable companies, silly!


There is this, but he is also a quasi-government figure. As Trump weakens he will reappear in that sphere.


you’re right, the bonds of this overlevered money losing enterprise AI and space cargo conglomerate should be trading at a premium


The bond price is an outcome of lenders judging, not a cause. Lenders have already judged.


Of course. But you have a bond buyer who can put their thumb on the scale, and has shown willingness to do so. That buyer also tried to interfere in (succeeded in interfering in) the business of law firms that worked for anti-Trump causes, has directly interfered in the renewable energy market to benefit its backers, etc.

What is to say that SpaceX and Oracle won't get these benefits? (Government buying bonds, trashing ratings agencies, leaning on banks to lend etc.)

Nothing obvious, I posit. So what is the value of a bond when a government is increasingly likely to manipulate the market for it?

And that is putting aside the second order of government interference: foreign governments putting their thumbs on the scale with their own investment funds and influenceable buyers, to buy influence over a government that favours these firms.


Nobody doubts all this. What the bond market is saying is that even with all of that, there are plenty of viable scenarios where bond holders aren't paid back.

After all, if you truly believe what you say with conviction, the sensible thing to do would be to buy up as many SpaceX bonds as you can if you think they're so undervalued.


Institutional bond traders are pretty sophisticated. They're aware of all the possibilities you mentioned, and pricing the bonds low anyway.


> Institutional bond traders are pretty sophisticated.

I am old enough to remember Long Term Capital Management.


I'm old enough to remember they weren't just buying and selling corporate bonds like most of the bond market. They were a hedge fund that used massive leverage to exploit tiny arbitrage opportunities between correlated securities.

There may be important factors that the bond market isn't aware of, just like most of them didn't know about the problems with mortgage-backed securities in 2008. But anything you and I are aware of, they probably are too.


> But anything you and I are aware of, they probably are too.

This is something that seems like it should be true but the subprime crisis (which didn't at all come out of the blue — anyone with any instinct at all should have understood when they first saw someone get given insane mortgages) argues against.


The subprime crisis was a story of perverse financial incentives causing the industry to play along with a fig-leaf statistical excuse for overrating derivative products, not bond raters botching ratings or missing "common-knowledge" info about specific mortgages.

If you want to suggest the subprime crisis as a mechanism for SpaceX bonds getting mispriced, you need to propose a model for how bond evaluators could be operating under a perverse incentive to under-rate it and somehow reap profits from doing so.


> If you want to suggest the subprime crisis as a mechanism for SpaceX bonds getting mispriced

I don't? I'm just observing that the bond market got something wrong — through an absolutely industry-wide blindness - that was indeed observed, years beforehand, by non-experts. Not my example, even. Just responding.


A lot of mortgages were trash but the securities were a new and complicated financial product with a plausible story. A few people figured out the problems, but Michael Burry for example did it by combing through thousands of pages of prospectuses.

SpaceX bonds are just plain ol' corporate bonds, the same stuff these investors have been analyzing for a very long time.


Max Keiser of Karmabanque was talking about and writing about it in public more than a year before Burry acted, with pretty solid predictions, without any obvious forensic action. He's quite bonkers, but if it was obvious to Keiser it should have been obvious to these highly intelligent people we're discussing.


Economists predict 500 out of every 10 recessions.


I too can play devil’s advocate

and now back to being an adult


Don't worry, they'll just build the data centers in NJ and still considered NY 1-20.

Sarcasm aside, I don't really know where they would build data centers in NYS. Electricity rate in northern and western NY is going thru the roof. ADK/Catskill have very sensitive environmental laws. Can't really build in lower hudson as real estate cost would be killer.


The Catskills have those environmental protection laws because they are the water source for NYC. It would be very stupid to relax those to build data centers.


NYC owns those lands for water. ADK has forever wild in NYS constitution. They are not gonna get relax for data center because they discharge water and noise and add significant infrastructure change.

Solar farm on the other hand might go up tho.


Replacing natural lands with solar farms is one of the stupidest things I've seen. They're doing that on some state parks near me.


Why are people disagreeing? I find it wild that people actually condone putting a solar farm on a state park. That land is supposed to be mostly natural and for public use. If the park wants to power itself using solar, then put it on the roofs and cover the parking areas with panels. Don't contract out a forest/field to be developed as a solar farm.

Maybe I'm jaded, but it seems like city people want to exploit the natural lands to produce their electricity. True NIMBY. It's hard to claim to care about the environment when being willing to destroy it when alternatives exist.


What would be unnatural lands that you'd find acceptable for use?


Existing infrastructure, such as roofs and parking lots.


Yeah why do I have to pay to get it in my roof when a solar company will buy land to put up solar cells. Just put it on my roof instead. I won't complain.


>Replacing natural lands with solar farms is one of the stupidest things I've seen

"Says here[1] on this study published by university lab funded by a consortium of the same interests that develop the solar that this is fine"

-smarmy HN linkposters

[1] <link to study that doesn't even support my thesis but you can't tell because it's behind a paywall>


> I don't really know where they would build data centers in NYS

They would build them as close to NYC as possible. Data Centers existed prior to AI boom. HFT, edge hosting, etc.


That's the insider joke. Look up NY1-NY4 data centers, they are all in NJ across the river. NYC just dump their shit into NJ is the usually move. But those areas are full now, and they don't really have anywhere to go but south jersey.


Go too far south and they start being Philadelphia data centers.


The AI ones are a hundred times larger. Formerly you needed a building to house racks where each customer had a few servers, now they want a megacomplex all dedicated to one purpose. Few of these are even getting completed, most are abandoned for cost reasons.


I've been reading Alastair Reynolds's books. They are pretty good.


OK. I've read it a few times and still don't understand. Where is the distributed part? You store data in a single transaction into postgres. What/who is notifying the message queue?


You build a distributed system on top of this! For example, you may have many distributed workers durably executing workflows from the Postgres-backed task queue. The Postgres transactions allow you to atomically perform operations spanning both your task queue and your business data.

Here's another blog post about how a Postgres-backed task queue can run at scale: https://www.dbos.dev/blog/making-postgres-queues-scale


I've been writing distributed workers for ages with stored functions that have a SELECT FOR UPDATE query.

When workers query the db for jobs the rows get locked by the select and there are no race conditions or duplicate assigned jobs


I'm wondering. Just wondering? Will they ever support multiple storage engines like MariaDB? Having a storage engine that support OLTP or OLAP or append-only would be cool. I totally understand if they don't want to do that.


MySQL has Galera cluster for that.


More accurately, MariaDB has Galera for that. MySQL Galera is EOL in a few months [1], which is understandable given the change in ownership.

[1] https://mariadb.com/resources/blog/upgrade-now-announcing-my...


And Group Replication


And percona xtradb cluster


I just came out of an initial interview for a rather senior enterprise position. It was bad. It was bad in a sense that they really didn't know what to do. The senior manager (EVP level person) asked me about LLM for code generation. I told them about my aws kiro-cli experience. They literally asked me to sit down and show them how they can do it. I'm pretty sure they want me to back for another round.

This whole thing reminds me of when I was in school, showing old timers who to use MS Office and VBA.


What's the difference between feather and parquet in terms of usage? I get the design philosophy, but how would you use them differently?


parquet is optimized for storage and compresses well (=> smaller files)

feather is optimized for fast reading


Given the cost of storage is getting cheaper, wouldn't most firms want to use feather for analytic performance? But everyone uses parquet.


You can, still, gain a lot of performance by doing less I/O.


There's definitely a "everyone uses it because everyone uses it" effect.

Feather might be a better fit for sime yse cases, but parquet has fantastic support and is still a pretty good choice for things that feather does.

Unless they're really focussed on eaking out every bit of read performance, people often opt for the well supported path instead.


What people have done in the face of cheaper storage is store more data.


Storage is cheap but bandwidth no.


Storage getting cheaper did not really reach the cloud providers and for self-hosting it has recently gotten even more expensive due to AI bs.


And now there's Lance! https://lance.org/



I read that. But afaik, feather format is stable now. Hence my confusion. I use parquet at work a lot, where we store a lot of time series financial data. We like it. Creating the Parquet data is a pain since it's not append-able.


Generally Parquet files are combined in an LSM style, compacting smaller files into larger ones. Parquet isn't really meant for the "journal" of level-0 append-one-record style storage, it's meant for the levels that follow.


So feather for journaling and parquet for long term processing?


I still don't understand what happened to using Apache Avro [1] for row-oriented fast write use cases.

I think by now a lot of people know you can write to Avro and compact to Parquet, and that is a key area of development. I'm not sure of a great solution yet.

Apache Iceberg tables can sit on top of Avro files as one of the storage engines/formats, in addition to Parquet or even the old ORC format.

Apache Hudi[2] was looking into HTAP capabilities - writing in row store, and compacting or merge on read into column store in the background so you can get the best of both worlds. I don't know where they've ended up.

[1] https://avro.apache.org/

[2] https://hudi.apache.org/


You basically can't do row by row appends to any columnar format stored in a single file. You could kludge around it by allocating arenas inside the file but that's still a huge write amplification, instead of writing a row in a single block you'd have to write a block per column.


You can do row by row appends to a Feather (Arrow IPC — the naming is confusing). It works fine. The main problem is that the per-append overhead is kind of silly — it costs over 300 bytes (IIRC) per append.

I wish there was an industry standard format, schema-compatible with Parquet, that was actually optimized for this use case.


Creating a new record batch for a single row is also a huge kludge leading to lot of write amplification. At that point, you're better off storing rows than pretending it's columnar.

I actually wrote a row storage format reusing Arrow data types (not Feather), just laying them out row-wise not columnar. Validity bits of the different columns collected into a shared per-row bitmap, fixed offsets within a record allow extracting any field in a zerocopy fashion. I store those in RocksDB, for now.

https://git.kantodb.com/kantodb/kantodb/src/branch/main/crat...

https://git.kantodb.com/kantodb/kantodb/src/branch/main/crat...

https://git.kantodb.com/kantodb/kantodb/src/branch/main/crat...


> Creating a new record batch for a single row is also a huge kludge leading to lot of write amplification.

Sure, except insofar as I didn’t want to pretend to be columnar. There just doesn’t seem to be something out there that met my (experimental) needs better. I wanted to stream out rows, event sourcing style, and snarf them up in batches in a separate process into Parquet. Using Feather like it’s a row store can do this.

> kantodb

Neat project. I would seriously consider using that in a project of mine, especially now that LLMs can help out with the exceedingly tedious parts. (The current stack is regrettable, but a prompt like “keep exactly the same queries but change the API from X to Y” is well within current capabilities.)


Frankly, RocksDB, SQLite or Postgres would be easy choices for that. (Fast) durable writes are actually a nasty problem with lots of little detail to get just right, or you end up with corrupted data on restart. For example, blocks may be written out of order so on a crash you may end up storing <old_data>12_4, and if you trust all content seen in the file, or even a footer in 4, you're screwed.

Speaking as a Rustafarian, there's some libraries out there that "just" implement a WAL, which is all you need, but they're nowhere near as battle-tested as the above.

Also, if KantoDB is not compatible with Postgres in something that isn't utterly stupid, it's automatically considered a bug or a missing feature (but I have plenty of those!). I refuse to do bug-for-bug compatible and there's some stuff that are just better not implement in this millennia, but the intent is to make it be I Can't Believe It's Not Postgres, and to run integration tests against actual everyday software.

Also, definitely don't use KantoDB for anything real yet. It's very early days.


> Frankly, RocksDB, SQLite or Postgres would be easy choices for that. (Fast) durable writes are actually a nasty problem with lots of little detail to get just right, or you end up with corrupted data on restart. For example, blocks may be written out of order so on a crash you may end up storing <old_data>12_4, and if you trust all content seen in the file, or even a footer in 4, you're screwed.

I have a WAL that works nicely. It surely has some issues on a crash if blocks are written out of order, but this doesn’t matter for my use case.

But none of those other choices actually do what I wanted without quite a bit of pain. First, unless I wire up some kind of CDC system or add extra schema complexity, I can stream in but I can’t stream out. But a byte or record stream streams natively. Second, I kind of like the Parquet schema system, and I wanted something compatible. (This was all an experiment. The production version is just a plain database. Insert is INSERT and queries go straight to the database. Performance and disk space management are not amazing, but it works.)

P.S. The KantoDB website says “I’ve wanted to … have meaningful tests that don’t have multi-gigabyte dependencies and runtime assumptions“. I have a very nice system using a ~100 line Python script that fires up a MySQL database using the distro mysqld, backed by a Unix socket, requiring zero setup or other complication. It’s mildly offensive that it takes mysqld multiple seconds to do this, but it works. I can run a whole bunch of copies in parallel, in the same Python process even, for a nice, parallelized reproducible testing environment. Every now and then I get in a small fight with AppArmor, but I invariably win the fight quickly without requiring any changes that need any privileges. This all predates Docker, too :). I’m sure I could rig up some snapshot system to get startup time down, but that would defeat some of the simplicity of the scheme.


And I have a system that launches Postgres in a container as part of a unit test (a little wrapper around https://crates.io/crates/pgtemp). It's much better than nothing, but the test using Postgres takes 0.5 seconds when the same business logic run against an in-memory implementation takes 0.005s.


Agreed.

There is room still for an open source HTAP storage format to be designed and built. :-)


Have you considered something like iceberg tables?


Yes, but parquet hates small files.


You can't compact? i.e. iceberg maintenance


We might be doing something wrong, but we saw significant performance degradation for both ingestion and query when doing compaction when it comes to finance data during trading hours.


Feather (Arrow IPC) is zero copy and an order of magnitude simpler. Parquet has a lot of compatibility issues between readers and writers.

Arrow is also directly usable as the application memory model. It’s pretty common to read Parquet into Arrow for transport.


When you say compatibility issues, you mean they are more problematic or less?

It’s pretty common to read Parquet into Arrow for transport.

I'm confused by this. Are you referring to Arrow Flight RPC? Or are you saying distributed analytic engine use arrow to transport parquet between queries?


Not the OP, but Parquet compatibility issues are usually due to the varying support of features across implementations. You have to take that into account when writing Parquet data (unless you go with the defaults which can be conservative and suboptimal).

Recently we have started documenting this to better inform choices: https://parquet.apache.org/docs/file-format/implementationst...


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: