Hacker Newsnew | past | comments | ask | show | jobs | submit | PeterCorless's commentslogin

This is actually a great article. We, as an industry, need to make sure that we address the points raised in this article. Especially What is, and what isn't a "streamhouse?"

How would you know you've achieved "streamhouse?" What objective requirements does it need to meet?

Disclosure: I work at Redpanda, which was part of the launch, and I will be hosting a webinar soon with some of the companies part of the technical working group. I'll try to raise some of the points made in this blog as part of the discussion if I have the chance.


This is really cool stuff. Amazing the efficiencies Pinot brings to data systems.

[Disclosure: I used to work at StarTree. Still a fan!]


I think that's the issue: a human can't. We'll need to set up "radar" agents to find out new tools for agents.


> Agents are cardinality-hungry. They want the high-cardinality data you'd normally drop: individual trace IDs, per-request attributes, full tag sets. They are very patient. They will sift through it.

The agents themselves are not likely going to be doing the high cardinality queries or they will keel over. They have limited memory buffers. They will take many seconds to return results. They are likely going to be limited in terms of QPS.

From the blog: > Apache Iceberg, with data stored as Parquet on S3, and most of the system implemented in Go

You have just ensured that queries will have a p99 >1 second. This is kind of antithetical to having an agent be fast.

You couldn't run any sort of real-time service, where hundreds of thousands to millions of events were occurring per second, and you needed to adjust to that in milliseconds.

The terms "p99" and "QPS" do not occur anywhere in the article. Which leaves the question of scalability to a user's imagination.

I applaud the direction. I am looking for objective evidence.


Author here - you are right that this architecture is not going to deliver super fast queries. But that's a tradeoff we're making: Agents don't need super fast queries to triage your software issues. In fact, the agents are extremely good at triage by fanning out to explore hypotheses against the telemetry. What they need is a datastore that allows them to run a _ton_ of queries in parallel, on the cheap. Data Lake architectures like ours provide exactly this.

At the end of the day, we're less focused on traditional database query metrics. We're optimizing for higher level outcomes, think mean time to remediation and such.

> The agents themselves are not likely going to be doing the high cardinality queries or they will keel over. They have limited memory buffers. They will take many seconds to return results. They are likely going to be limited in terms of QPS.

You're right that if you simply give an LLM a tool to query a massive high cardinality dataset, it's going to blow itself (its context window) up. That's not what we're doing: instead we harness the llm with purpose-built tools + prompt + context + other engineering to ensure the agent can explore the data and make progress, even if it does run a dumb query on occasion.


Vera does what NVIDIA calls Spatial Multithreading, "physically partitioning each core’s resources rather than time slicing them, allowing the system to optimize for performance or density at runtime." A kind of static hyperthreading; you get two threads per core.

It's somewhat different from how x86 chips do simultaneous multithreading (SMT),


Seems like curious terminology from NV. In estabilished use, SMT means executing instructions from several cpu threads concurrently in the OOO CPU's execution units so they are not starved from work, whereas timeslicing conventionally means context switching between threads/processes, alternating temporally.

In operating systems timeslicing means giving a quantum of execution time to each process, and context switching between processes. Not normally a term used in computer architecture but possibly the characterisation would fit a barrer processor rather than SMT.


This is the related benchmark blog from Redpanda [disclosure: I work for Redpanda and I helped write this. Credit to Travis Downs & others at Redpanda for the heavy lifting on the testing and analysis.]

https://www.redpanda.com/blog/nvidia-vera-cpu-performance-be...


The reason being? IP proxy gateways. They obviated the need to move away from the limited address space of IPv4. Which was 90% of the reason to do IPv6.


Should be far, far larger news, to be honest.


So much that we presume in the modern cloud wasn't a given when Apache Kafka was first released in 2011.

kevstev wrote just above about Kafka being written to run on spinning disks (HDDs), while Redpanda was written to take advantage of the latest hardware (local NVMe SSDs). He has some great insights.

As well, Apache Kafka was written in Java, back in an era when you were weren't quite sure what operating system you might be running on. For example, when Azure first launched they had a Windows NT-based system called Windows Azure. Most everyone else had already decided to roll Linux. Microsoft refused to budge on Linux until 2014, and didn't release its own Azure Linux until 2020.

Once everyone decided to roll Linux, the "write once run everywhere" promise of Java was obviated. But because you were still locked into a Java Virtual Machine (JVM) your application couldn't optimize itself to the underlying hardware and operating system you were running on.

Redpanda, for example, is written in C++ on top of the Seastar framework (seastar.io). The same framework at the heart of ScyllaDB. This engine is a thread-per-core shared-nothing architecture that allows Redpanda to optimize performance for hardware utilization in ways that a Java app can only dream of. CPU utilization, memory usage, IO throughput. It's all just better performance on Redpanda.

It means that you're actually getting better utility out of the servers you deploy. Less wasted / fallow CPU cycles — so better price-performance. Faster writes. Lower p99 latencies. It's just... better.

Now, I am biased. I work at Redpanda now. But I've been a big fan of Kafka since 2015. I am still bullish on data streaming. I just think that Apache Kafka, as a Java-based platform, needs some serious rearchitecture,

Even Confluent doesn't use vanilla Kafka. They rewrote their own engine, Kora. They claim it is 10x faster. Or 30x faster. Depending on what you're measuring.

1. https://www.confluent.io/confluent-cloud/kora/

2. https://www.confluent.io/blog/10x-apache-kafka-elasticity/


There have been annual layoffs at RedHat since 2023. This year they just laid off more. The layoffs this year are expected to be "a low single digit percentage of our global workforce." Which will likely include hundreds of folks at Red Hat.

1. https://www.cio.com/article/4084855/ibm-to-cut-thousands-of-...

2. https://www.newsobserver.com/news/business/article312796900....


There haven't been layoffs at Red Hat after 2023, whereas according to your statement there should have been two more rounds. The layoffs from your articles are at IBM, and did not affect Red Hat.


Thank you for the correction.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: