Open-sourcing a 10x reduction in Apache Cassandra tail latency (opens in new tab)

(engineering.instagram.com)

408 pointsmikeyk8y ago164 comments

164 comments

> The graph shows that a Cassandra server instance could spend 2.5% of runtime on garbage collections instead of serving client requests. The GC overhead obviously had a big impact on our P99 latency

No, this is not obvious. If you have a fully concurrent GC then spending 25 out of 1000 CPU cycles on memory management does not "obviously" have an impact on your 99th percentile latency. It would primarily impact your throughput (by 2.5%), just like any other thing consuming CPU cycles.

> We defined a metric called GC stall percentage to measure the percentage of time a Cassandra server was doing stop-the-world GC (Young Gen GC) and could not serve client requests.

Again, this metric doesn't tell you anything if you don't know how long each of the pauses are. If they are at the limit infinitesimally small then you are again only measuring the impact on throughput, not latency.

Certainly, GCs with long STW pauses do impact latency, but then you need to measure histograms of absolute pause times, not averages of ratios relative to application time. That's just a silly metric.

And neither does the article mention which JVM or GC they're using. Absent further information they might have gotten their 10x improvement relative to some especially poor choice of JVM and GC.

_ivvf8y ago

you clearly didn't read the post very closely. They said 2.5% of CPU cycles were spent on stop-the-world young generation collections, not on the sum total of all memory mangement. That means that 2.5% of the time the app is entirely stalled on just these collections. Given that stop-the-world pauses are never evenly distributed throughout time, it should be very much expected that this much GC stalling would affect p99 latencies.

It's pretty much accepted everywhere that GCs perform terribly for databases. Modern GCs are great at handling small, very short-lived memory allocations, and that's about it. Just about any other workload and manual memory management ends up being a much better use of your time than GC tuning.

b4lancesh33t8y ago

> Given that stop-the-world pauses are never evenly distributed throughout time

That is not a given. And, even distribution is only part of the equation. If they are sufficiently short, then even being somewhat unevenly distributed should not have much of an impact on latency. For example, if the max length of a pause were 1ms, and 99p latency were 15ms, you'd have to be fairly unlucky to see a 33% increase in latency99 due to GC. That would entail 5 of 25 pauses happening during a 20ms period in a 1s window.

(This idea is not purely hypothetical. For example, Go's GC has very low STW periods.)

> It's pretty much accepted everywhere

Eh. Apparently everyone thinks C is the best language for cryptography and other secure but not particularly perf sensitive code. Go figure. Sometimes the wisdom of the masses is not wisdom. Best not to appeal to it during argumentation.

3 more replies

nvarsj8y ago

So why do people keep building latency sensitive things in the JVM? And then they manage to get hugely popular?

Cassandra is a constant struggle with the GC. I’d guess the cost of running it is at least an order of magnitude greater compared to if it had been implemented in c++ or something more sensible.

5 more replies

foolfoolz8y ago

classic hacker news comment.

this thing you built and open sourced, has gotten you real measurable results? allow me to list the many ways you’re probably wrong and doing it incorrectly

discoursism8y ago

Measurable results are all well and good, but it can be helpful to know how the baseline was established. Measurable results aren't "portable" without a well-established baseline.

fdeliege8y ago

Both code and benchmark are open sourced. We'd love to hear how it performs for you.

viraptor8y ago

This is a valid criticism of the methodology / explanation. It's not about the results. You can agree with the positive results (and they're great! - you've done awesome work and clearly show an improvement) and still say the explanation how/why they were achieved is not great.

1 more reply

teacpde8y ago

> If you have a fully concurrent GC then spending 25 out of 1000 CPU cycles on memory management does not "obviously" have an impact on your 99th percentile latency.

I try to understand the meaning. Is it saying the latency caused be GC is applied to all requests, not just the ones that observe 99th percentile latency?

dtparr8y ago

It's saying that whether it affects latency or just throughput depends on how those pauses are distributed in absolute terms, not just the ratio. There's a big difference in 99th percentile latency between a 1ms pause every 400ms and a 10 second pause every 67 minutes, but they both work out to 2.5% by the ratio metric.

So yes, at the `infinitesimally small` end, time would be 'stolen' evenly from all request threads and would not be a contributing factor to the 99th percentile.

the84728y ago

No, that would be an incremental GC working in very small time slices.

A concurrent GC spends CPU cycles on different cores to do its work, which means it will not cause latency outliers in the threads processing the requests. They are still CPU cycles you don't have to serve other requests, hence they still affect throughput.

That is a simplified explanation of course, there are a lot of caveats.

In my original post I was mostly speaking about the measurement though, since they are measuring throughput when they are concerned about latency, those are somewhat related but depending on circumstances only weakly so.

2 more replies

dikanggu8y ago

We do want to contribute our work back to the Cassandra upstream, instead of keeping it as a fork. So that more users from C* community can benefit from the improvements. The pluggable storage engine is an ambitious project (https://issues.apache.org/jira/browse/CASSANDRA-13474). Any help will be appreciated!

russellspitzer8y ago

Saw you talking about this on the Distributed Data Show

https://academy.datastax.com/content/distributed-data-show-e...

gfosco8y ago

RocksDB is used all over Facebook, powers the entire social graph. Great storage engine that pairs well with multiple DBMS: MySQL, Mongo, Cassandra... We'll be at Percona Live 2018 in April, giving several talks, and are looking forward to hanging out and talking with users in our lounge area. We're working hard to support our open source community as well! https://github.com/facebook/rocksdb

openasocket8y ago

I'm not an expert on these things, but it seems to me if you're implementing a database in Java you wouldn't want to keep your data on the JVM Heap, as this seems to indicate. My understanding is that in most applications (like servers) the average object lives for a very short period of time, and most GC implementations are built from that idea. But, in a database, especially an in-memory database, the majority of the objects are going to live for a very long time. That makes the mark phase of GC a lot more expensive, puts more pressure on the generations, etc.

Is my guess here correct, or are there things I'm missing or mistaken on?

jakewins8y ago

This is correct; the standard approach here is to use regular c-style memory management for the data the system is managing, and the JVM heap only for the database "infrastructure".

This hybrid approach gives the benefit of a managed runtime and safety of GC for most of your code, but allows the performance of raw pointers/malloc for key code paths.

Some examples of this pattern on the JVM:

- The Neo4j Page Cache, Muninn, https://github.com/neo4j/neo4j/blob/3.4/community/io/src/mai...

- The Netty projects implementation of jemalloc for the JVM: https://github.com/netty/netty/blob/4.1/buffer/src/main/java...

ndesaulniers8y ago

> but allows the performance of ... malloc for key code paths.

Everything is relative, I guess.

1 more reply

tkahnoski8y ago

If so any JVM based datastore could probably benefit.

I wonder how long before we see a similar result from ElasticSearch. (Only other huge JVM based store I can think of).

2 more replies

jjirsa8y ago

Cassandra does a bunch of stuff off-heap - we keep things like Bloom Filters, our compression offsets (to seek into compressed data files), and even some of the memtable (the in-memory buffer before flushing) in direct memory, primary for the reasons you describe.

We still have "other" things on-heap. The biggest contributor to GC pain tends to be the number of objects allocated on the read path, so this patch works around that by pushing much of that logic to rocksdb.

There are certainly other things you can do in the code itself that would also help - one of the biggest contributors to garbage is the column index. CASSANDRA-9754 fixes much of that (jira is inactive, but the development work on it is ongoing).

wonnage8y ago

The purpose of separating into young and old generation is that it's easier to find dead objects in the young generation (as you said, average object lives for a short period of time). You only have to scan this subset for a minor GC. It doesn't really matter how many long-lived objects you have as long as you can avoid needing to do a major GC.

openasocket8y ago

Don't you still need to scan the old generation during minor GC, in case a field in one of the older objects was modified to point to an object in the young generation? Or are there optimizations you can use to quickly and efficiently find references from the older generation to the younger?

1 more reply

coredog648y ago

For a long time, the guidance was to install jemalloc and then use off-heap objects. I can’t recall what it was, but that broke in the 3.0.x series and is unlikely to return. The feature stream (3.1.x) allegedly can use jemalloc again, but we’ve been slow to adopt it so I can’t provide proof.

ADefenestrator8y ago

The good news is, at least for my workloads, garbage collection is significantly better in 3.0 than 2.1 even without off-heap objects. Not sure if it's generating less or just generating it in a way that's easier to collect cheaply, but I saw pause times and total collection work drop significantly with the same settings (G1 collector)

StreamBright8y ago

If you want to keep data off heap you need to use sun.misc.Unsafe and allocate / free by yourself. I guess it is called unsafe for a reason. With G1GC you can do magical things to reduce the GC overhead which I always recommend as the first step before trying off heap.

kodablah8y ago

ByteBuffer.allocateDirect is another off-JVM-heap solution that's not marked unsafe.

the84728y ago

You don't even have to go through Unsafe. ByteBuffer.allocateDirect() gives you a chunk of off-heap memory.

haglin8y ago

"To reduce the GC impact from the storage engine, we considered different approaches and ultimately decided to develop a C++ storage engine to replace existing ones."

I wonder how the numbers would have looked with the new low latency GC for Hotspot (ZGC). https://wiki.openjdk.java.net/display/zgc/Main

Early results from SPECjbb2015 are impressive. https://youtu.be/tShc0dyFtgw?t=5m1s

tibbetts8y ago

Yes, also Azul Zing. Really anytime someone says they have a problem with GC and suggests spending a million dollars of engineer time building a new system, they should consider Zing first. It works and is a way more efficient way of spending money to fix GC latency problems.

ADefenestrator8y ago

For a small to medium sized shop, sure. For someplace with thousands or tens of thousands of nodes, the new system ends up cheaper in the long run.

majidazimi8y ago

Because GC related issues don't undergo from "problem" state to "solved" state. It is just a never ending stream of issues, that the team need to resolve, specially in a database realm in which metrics are hugely workload dependent.

itronitron8y ago

yes, a comparison across multiple JVMs would be nice

Thaxll8y ago

Weird, did they try to use https://www.scylladb.com/?

jjirsa8y ago

Why throw away something proven to run at massive scale, that you understand and trust for something that's new, has never been run at that scale, and you have no experience running? If you have a team of software engineers, and the latency problem is a software problem, fix the software problem.

When you already know Cassandra, and you already know RocksDB, and you already have an engineering team, it makes far more sense to combine the two things you know how to use at scale than to try to use some new thing NOBODY has run at scale.

welder8y ago

> some new thing NOBODY has run at scale

Outbrain uses ScyllaDB in production at scale across multiple data centers. Not sure if it's Instagram scale, but still enough to prove it's reliability and performance.

https://www.outbrain.com/techblog/2016/08/scylladb-poc-not-s...

1 more reply

en4bz8y ago

I was going to say the same thing. It seems pretty clear at this point that Java is not a good programming language to build a database on if you care about strong 99% latency guarantees. The engineers in the article came to this conclusion and so did the Scylla people years ago.

Scylla is AGPL for the OSS version though so testing it out would not be an option without getting a commercial license first.

geofft8y ago

> Scylla is AGPL for the OSS version though so testing it out would not be an option without getting a commercial license first.

Huh? The AGPL is not a non-commercial-use-only license.

If you have proprietary software that you would like to combine with AGPL code (i.e., not interact with as a service) and is available to the general public over the Internet, and you want keep your code proprietary, sure, you may not want to use the AGPL. But you could say the same thing about proprietary software you want to combine with GPL code and sell to the general public.

If you're either using the software through it's existing defined public interfaces, or you're okay releasing anything you modify or link into the software, the AGPL (and the GPL) are fine. Lots of people distribute proprietary products that include GPL code, like Chromebooks, Android phones, routers, GitHub Enterprise, etc. We figured out years ago that the Linux kernel is not just a non-commercial product. Why are we having the same misconceptions about the AGPL?

2 more replies

ghshephard8y ago

There are those who've deployed on Java with tight latency requirements: https://martinfowler.com/articles/lmax.html?t=1319912579 - Benchmarked at around 6 million transactions/second.

The issue isn't so much Java the language, as it is being aware of the GC, and developing with it in mind.

1 more reply

glommer8y ago

(ScyllaDB employee here)

I don't believe one would need a commercial license just to test a product in any way? They are not making that part of any product at that point, so no concerns here.

1 more reply

gnud8y ago

The server is AGPL. The client is Apache licensed. So I don't see a problem with using the AGPL version in commercial product.

Noone claims that a product using the MySql driver is a derivative work of the MySql server?

Edit: Of course, IANAL...

1 more reply

HippoBaro8y ago

Doesn't AGPL allow commercial use?

3 more replies

jakelarkin8y ago

- has anyone run it FB scale? for how long?

- how many experienced scylladb devops are there globally that we can hire?

Those questions asked at BigTechCo before it adopts somebody elses tech.

FB already operates RocksDb and Cassandra so there's way less technical, career, financial risk for just hacking the two together with some aggressive refactoring.

polskibus8y ago

Does FB still use Cassandra? I thought they abandoned them ages ago and then databricks picked it up?

3 more replies

HippoBaro8y ago

My thought exactly. Would be interesting to know if they did and if yes, why they chose to develop something in-house anyway.

tschellenbach8y ago

For Stream's feed tech we also moved from Cassandra to an in-house solution on top of RocksDB. It's been a massive performance and maintenance improvement. This StackShare explains how Stream's stack works. It's based on Go, RocksDB and Raft: https://stackshare.io/stream/stream-and-go-news-feeds-for-ov...

3uclid8y ago

Unrelated: as a CS undergrad, I read this article and was immediately inspired. This is definitely the type of work I want to be doing when I graduate (infrastructure engineering). But my next thought was: where do I start?!

Any advice?

en4bz8y ago

CMU Database Group Lectures: https://www.youtube.com/channel/UCHnBsf2rH-K7pn09rb3qvkA

therealdrag08y ago

I'd say no matter what kind of job you get, you can put 10% of your time into similar problems. Even simple CRUD apps can have interesting problems like this. In my experience every project has instances of engineers shooting themselves in the foot, or unforeseen problems cropping up. If you have a bit of self-motivation you can dig into them and learn a lot and improve things. I do this and find it very satisfying.

jjirsa8y ago

Happy to help you get started working on Cassandra. http://cassandra.apache.org/doc/latest/development/patches.h... Has some basic entry pointers. There’s also a dev mailing list that’s reasonable active.

ddorian438y ago

Still in school ? (don't understand different <type>grad). See: GSOC Seastar Framework https://summerofcode.withgoogle.com/organizations/6190282903...

3uclid8y ago

Yeah, still in school (3rd year). I have intern experience, but it seems like these type of positions are way too advanced for me at the moment. Just unsure how to progress...

1 more reply

StreamBright8y ago

In a similar situation we just adjust the GC and started to use G1GC which resulted in similar numbers.

coryfoo8y ago

I bet that didn't take N engineers 12 months to build out, either

threeseed8y ago

Cassandra uses G1GC by default.

If it was as simple as tweaking a few GC settings to get 10x improvement pretty sure Datastax would've done it by now.

2 more replies

StreamBright8y ago

2 engineers, 2 weeks because we had to evaluate every change we made with production traffic.

fdeliege8y ago

Join our meetup to chat with some of the developers: https://www.meetup.com/Apache-Cassandra-Bay-Area/events/2483...

jjirsa8y ago

So sad I’m not in town that week

en4bz8y ago

Has any tried running Casandra on Azul Zing[1]? The slowdown here is not surprisingly related to GC pauses which Azul has eliminated in Zing.

[1] https://www.azul.com/products/zing/

rbranson8y ago

The licensing cost of Zing generally makes this a bad trade-off. It's much cheaper to purchase more hardware. Zing is targeted at vertically scaling very large JVM heaps, where it's valuable to have massive amounts of data on a single, big machine.

nitsanw8y ago

As an ex-Azul employee I can say there's a good number of Azul clients using a Zing+Cassandra setup, so the price point is right for some people at the very least. Zing licence cost has also changed in recent years (3.5k per server last I looked, and that is before you haggle some bulk deal) so not sure if your impression is calibrated to that new price point.

jjirsa8y ago

Have friends who have used it, they report that it works reasonably well. Especially in p99.

truth_seeker8y ago

By what factor/magnitude p99 was improved ? Any idea ?

spockz8y ago

Actually, it appears that is one of the premises[1] they sell Zing on.

[1]: https://www.azul.com/solutions/cassandra/

adrianratnapala8y ago

As a Java scoffer trying to be fair-minded, I resisted the urge to joke that "it's was the GC, stupid" and assume that a big project like Cassandra had somehow worked around the GC latency problems.

But, what? It turns out the article is really about replacing Java with C++.

cestith8y ago

It's about using something in one language for its features and only porting the critical sections to C++ via a clean API. This is the sort of advice we've been giving people for decades. Choose the language for what you want to build, measure and profile performance if necessary, find the bottleneck on the hot path, decouple that from the bulk of the code, and reach to a lower level for performance only in that clearly defined section.

They managed to generalize one application that meets their feature needs to be a front end to another existing application with fewer features but better performance as a back end. They're optimizing their hot path by decoupling it from the rest of the application and handing off to C++ code they didn't even have to write. Adding pluggable storage engines to Cassandra means that if they make the API smooth enough they can have engines in C, C++, Erlang, Go, Rust, ML, or whatever in the future without changing their front end. That's a big win even beyond this tail latency issue.

majidazimi8y ago

Well, other than storage engine, the next big part of a database software is the query planner/optimizer which Cassandra doesn't have (due to simple KV nature of it). So there isn't much remaining. In a long term plan, rewrite them all and you have single code base and you'll benefit from mighty C++ in other components of the database. And there is still room for more optimizations: SIMD, ...

The GC problem is not limited to C*. This shit(virtual machine) is hitting the whole Hadoop stack: HDFS, Hive, Spark, Flink, Pig...

Immense number of tickets in any fairly large cluster is related somewhat to GC and JVM behavior.

cmrdporcupine8y ago

I remember using quite early versions of Cassandra back in an ad-tech startup I was at back in 2009 or 2010, spending unfortunate amounts of time fighting the JVM GC and trying to tune things so it behaved responsibly. It was a real problem then and I know a lot of work went into fixing GC behaviour. Then I stopped using Cassandra for work, but it's unfortunate this is still an issue?

What I took out of that is that I really feel like something like Cassandra is better suited to implementation in a language like C++ or Rust. And I believe others have since come along and done this.

I really liked the gossip-based federation in Cassandra though.

ADefenestrator8y ago

It's still an issue, but a lot less of one. 3.0 is a big improvement in terms of GC behavior. Haven't tested 3.11.x yet, but it look in theory like a decent improvement in terms of rounding off corner cases and adding instrumentation.

estebank8y ago

It sounds like you might be interested in TiKV.

https://github.com/pingcap/tikv

cmrdporcupine8y ago

Thanks.

Since coming to Google I don't get the opportunity to compare/evaluate/deploy tools like this anymore. Smarter people than me make choices like that :-)

bfrog8y ago

Meanwhile scylladb looks like a better option for numerous reasons

yazr8y ago

Or just try and benchmark Azul VM with pause-less GCs ?!

(I have used Azul in low-latency production environments. It has pros and cons but it certainly beats re-writing the storage layer... )

truth_seeker8y ago

Curious to know the cons of using it, except being commercial.

manigandham8y ago

> except being commercial

That's the biggest, especially for when it's a Facebook company. Otherwise it works well but can be pricey.

The JVM is getting a new fully concurrent collector though called Shenandoah: https://www.google.com/search?q=shenandoah+gc

1 more reply

yazr8y ago

Needs a stronger machine to be effective (more cores & more memory)

Minor configuration issues (we had a very complex environment, custom kernel, weird network stuff, JNIs)

jjirsa8y ago

Nicely done! Looking forward to the pluggable storage engine.

pas8y ago

The JIRA tickets don't really shine with much hope :/

https://issues.apache.org/jira/browse/CASSANDRA-13474 [2 comments from 2017 Apr] https://issues.apache.org/jira/browse/CASSANDRA-13475 [~100 comments, but the last one is from 2017 Nov, by the InstaG engineer]

And the Rocksandra fork is already ~3500 commits behind master, so upstreaming this will be interesting.

Oh, and the Rocksandra fork is already kind of abandoned - no commits since 2017 Dec. (which probably means this is not actually the code that runs under Instagram.)

dikanggu8y ago

This is the rocksandra branch, https://github.com/Instagram/cassandra/tree/rocks_3.0, we develop it on top of Cassandra 3.0. It's the code we are running on our production servers.

1 more reply

jjirsa8y ago

I'm a committer, I'm familiar with the JIRA ticket.

1 more reply

kiril-me8y ago

It would be great. But I don't think it could happen. The pluggable storage engine would greatly increase the cognitive complexity of the code.

jjirsa8y ago

Of course it could happen. Pluggable engine has a lot of positives - not only does it enable features like this, it also helps modularize the codebase making it more testable, and there are other people who will develop storage engines for their own use case over time (look at the evolution of - for example - MySQL storage backends for examples of this).

rbranson8y ago

Did you all find that there were changes to the Java heap/GC configuration that would make tuning this setup different? I imagine if most everything that "sticks" is moved off heap, the GC could be tuned more heavily for young gen throughput vs trying to balance it with long-lived objects.

dikanggu8y ago

Yeah, for Rocksandra, we are able to use much smaller heap size, and most of the objects are recycled during the young gen GC.

agnivade8y ago

> We also observed that the GC stalls on that cluster dropped from 2.5% to 0.3%, which was a 10X reduction!

Umm .. shouldn't the stalls go to 0, because now you have moved to C++ ? Or is this the time it takes for the manual garbage collection to occur ?

steeve8y ago

Why not use ScyllaDB ? (Serious)

cnlwsu8y ago

Answered a bit before but they have a team that knows c* well. Cassandra is proven to handle petabytes at scale in production systems.

xuanyue8y ago

Is there any trade off after replacing LSM tree-based storage engine to RocksDB storage engine?

irfansharif8y ago

RocksDB is also an LSM structured KV store.

welder8y ago

Great, now can you fix the Python Cassandra Driver to work in a multi-threaded application environment without the connection pooling bugs and default synchronous app-blocking (vs lazy-init) connection setup?

https://github.com/datastax/python-driver

ismail8y ago

So question:

Any thoughts on replacing HDFS + Yarn + Hive + HBASE with GulsterFS + Kubernetes + Cassandra

ddorian438y ago

Hbase is sync+globally sorted, while cassandra is not, so probably not.

alsadi8y ago

Can we add lz4 to the blend to reduce disk IO?

j / k navigate · click thread line to collapse

164 comments

the84728y ago

> We defined a metric called GC stall percentage to measure the percentage of time a Cassandra server was doing stop-the-world GC (Young Gen GC) and could not serve client requests.

And neither does the article mention which JVM or GC they're using. Absent further information they might have gotten their 10x improvement relative to some especially poor choice of JVM and GC.

_ivvf8y ago

b4lancesh33t8y ago

> Given that stop-the-world pauses are never evenly distributed throughout time

(This idea is not purely hypothetical. For example, Go's GC has very low STW periods.)

> It's pretty much accepted everywhere

3 more replies

nvarsj8y ago

So why do people keep building latency sensitive things in the JVM? And then they manage to get hugely popular?

Cassandra is a constant struggle with the GC. I’d guess the cost of running it is at least an order of magnitude greater compared to if it had been implemented in c++ or something more sensible.

5 more replies

foolfoolz8y ago

classic hacker news comment.

this thing you built and open sourced, has gotten you real measurable results? allow me to list the many ways you’re probably wrong and doing it incorrectly

discoursism8y ago

Measurable results are all well and good, but it can be helpful to know how the baseline was established. Measurable results aren't "portable" without a well-established baseline.

fdeliege8y ago

Both code and benchmark are open sourced. We'd love to hear how it performs for you.

viraptor8y ago

1 more reply

teacpde8y ago

> If you have a fully concurrent GC then spending 25 out of 1000 CPU cycles on memory management does not "obviously" have an impact on your 99th percentile latency.

I try to understand the meaning. Is it saying the latency caused be GC is applied to all requests, not just the ones that observe 99th percentile latency?

dtparr8y ago

So yes, at the `infinitesimally small` end, time would be 'stolen' evenly from all request threads and would not be a contributing factor to the 99th percentile.

the84728y ago

No, that would be an incremental GC working in very small time slices.

That is a simplified explanation of course, there are a lot of caveats.

2 more replies

dikanggu8y ago

russellspitzer8y ago

Saw you talking about this on the Distributed Data Show

https://academy.datastax.com/content/distributed-data-show-e...

gfosco8y ago

openasocket8y ago

Is my guess here correct, or are there things I'm missing or mistaken on?

jakewins8y ago

This is correct; the standard approach here is to use regular c-style memory management for the data the system is managing, and the JVM heap only for the database "infrastructure".

This hybrid approach gives the benefit of a managed runtime and safety of GC for most of your code, but allows the performance of raw pointers/malloc for key code paths.

Some examples of this pattern on the JVM:

- The Neo4j Page Cache, Muninn, https://github.com/neo4j/neo4j/blob/3.4/community/io/src/mai...

- The Netty projects implementation of jemalloc for the JVM: https://github.com/netty/netty/blob/4.1/buffer/src/main/java...

ndesaulniers8y ago

> but allows the performance of ... malloc for key code paths.

Everything is relative, I guess.

1 more reply

tkahnoski8y ago

If so any JVM based datastore could probably benefit.

I wonder how long before we see a similar result from ElasticSearch. (Only other huge JVM based store I can think of).

2 more replies

jjirsa8y ago

wonnage8y ago

openasocket8y ago

1 more reply

coredog648y ago

ADefenestrator8y ago

StreamBright8y ago

kodablah8y ago

ByteBuffer.allocateDirect is another off-JVM-heap solution that's not marked unsafe.

the84728y ago

You don't even have to go through Unsafe. ByteBuffer.allocateDirect() gives you a chunk of off-heap memory.

haglin8y ago

"To reduce the GC impact from the storage engine, we considered different approaches and ultimately decided to develop a C++ storage engine to replace existing ones."

I wonder how the numbers would have looked with the new low latency GC for Hotspot (ZGC). https://wiki.openjdk.java.net/display/zgc/Main

Early results from SPECjbb2015 are impressive. https://youtu.be/tShc0dyFtgw?t=5m1s

tibbetts8y ago

ADefenestrator8y ago

For a small to medium sized shop, sure. For someplace with thousands or tens of thousands of nodes, the new system ends up cheaper in the long run.

majidazimi8y ago

itronitron8y ago

yes, a comparison across multiple JVMs would be nice

Thaxll8y ago

Weird, did they try to use https://www.scylladb.com/?

jjirsa8y ago

welder8y ago

> some new thing NOBODY has run at scale

Outbrain uses ScyllaDB in production at scale across multiple data centers. Not sure if it's Instagram scale, but still enough to prove it's reliability and performance.

https://www.outbrain.com/techblog/2016/08/scylladb-poc-not-s...

1 more reply

en4bz8y ago

Scylla is AGPL for the OSS version though so testing it out would not be an option without getting a commercial license first.

geofft8y ago

> Scylla is AGPL for the OSS version though so testing it out would not be an option without getting a commercial license first.

Huh? The AGPL is not a non-commercial-use-only license.

2 more replies

ghshephard8y ago

There are those who've deployed on Java with tight latency requirements: https://martinfowler.com/articles/lmax.html?t=1319912579 - Benchmarked at around 6 million transactions/second.

The issue isn't so much Java the language, as it is being aware of the GC, and developing with it in mind.

1 more reply

glommer8y ago

(ScyllaDB employee here)

I don't believe one would need a commercial license just to test a product in any way? They are not making that part of any product at that point, so no concerns here.

1 more reply

gnud8y ago

The server is AGPL. The client is Apache licensed. So I don't see a problem with using the AGPL version in commercial product.

Noone claims that a product using the MySql driver is a derivative work of the MySql server?

Edit: Of course, IANAL...

1 more reply

HippoBaro8y ago

Doesn't AGPL allow commercial use?

3 more replies

jakelarkin8y ago

- has anyone run it FB scale? for how long?

- how many experienced scylladb devops are there globally that we can hire?

Those questions asked at BigTechCo before it adopts somebody elses tech.

FB already operates RocksDb and Cassandra so there's way less technical, career, financial risk for just hacking the two together with some aggressive refactoring.

polskibus8y ago

Does FB still use Cassandra? I thought they abandoned them ages ago and then databricks picked it up?

3 more replies

HippoBaro8y ago

My thought exactly. Would be interesting to know if they did and if yes, why they chose to develop something in-house anyway.

tschellenbach8y ago

3uclid8y ago

Any advice?

en4bz8y ago

CMU Database Group Lectures: https://www.youtube.com/channel/UCHnBsf2rH-K7pn09rb3qvkA

therealdrag08y ago

jjirsa8y ago

ddorian438y ago

Still in school ? (don't understand different <type>grad). See: GSOC Seastar Framework https://summerofcode.withgoogle.com/organizations/6190282903...

3uclid8y ago

Yeah, still in school (3rd year). I have intern experience, but it seems like these type of positions are way too advanced for me at the moment. Just unsure how to progress...

1 more reply

StreamBright8y ago

In a similar situation we just adjust the GC and started to use G1GC which resulted in similar numbers.

coryfoo8y ago

I bet that didn't take N engineers 12 months to build out, either

threeseed8y ago

Cassandra uses G1GC by default.

If it was as simple as tweaking a few GC settings to get 10x improvement pretty sure Datastax would've done it by now.

2 more replies

StreamBright8y ago

2 engineers, 2 weeks because we had to evaluate every change we made with production traffic.

fdeliege8y ago

Join our meetup to chat with some of the developers: https://www.meetup.com/Apache-Cassandra-Bay-Area/events/2483...

jjirsa8y ago

So sad I’m not in town that week

en4bz8y ago

Has any tried running Casandra on Azul Zing[1]? The slowdown here is not surprisingly related to GC pauses which Azul has eliminated in Zing.

[1] https://www.azul.com/products/zing/

rbranson8y ago

nitsanw8y ago

jjirsa8y ago

Have friends who have used it, they report that it works reasonably well. Especially in p99.

truth_seeker8y ago

By what factor/magnitude p99 was improved ? Any idea ?

spockz8y ago

Actually, it appears that is one of the premises[1] they sell Zing on.

[1]: https://www.azul.com/solutions/cassandra/

adrianratnapala8y ago

As a Java scoffer trying to be fair-minded, I resisted the urge to joke that "it's was the GC, stupid" and assume that a big project like Cassandra had somehow worked around the GC latency problems.

But, what? It turns out the article is really about replacing Java with C++.

cestith8y ago

majidazimi8y ago

The GC problem is not limited to C*. This shit(virtual machine) is hitting the whole Hadoop stack: HDFS, Hive, Spark, Flink, Pig...

Immense number of tickets in any fairly large cluster is related somewhat to GC and JVM behavior.

cmrdporcupine8y ago

I really liked the gossip-based federation in Cassandra though.

ADefenestrator8y ago

estebank8y ago

It sounds like you might be interested in TiKV.

https://github.com/pingcap/tikv

cmrdporcupine8y ago

Thanks.

Since coming to Google I don't get the opportunity to compare/evaluate/deploy tools like this anymore. Smarter people than me make choices like that :-)

bfrog8y ago

Meanwhile scylladb looks like a better option for numerous reasons

yazr8y ago

Or just try and benchmark Azul VM with pause-less GCs ?!

(I have used Azul in low-latency production environments. It has pros and cons but it certainly beats re-writing the storage layer... )

truth_seeker8y ago

Curious to know the cons of using it, except being commercial.

manigandham8y ago

> except being commercial

That's the biggest, especially for when it's a Facebook company. Otherwise it works well but can be pricey.

The JVM is getting a new fully concurrent collector though called Shenandoah: https://www.google.com/search?q=shenandoah+gc

1 more reply

yazr8y ago

Needs a stronger machine to be effective (more cores & more memory)

Minor configuration issues (we had a very complex environment, custom kernel, weird network stuff, JNIs)

jjirsa8y ago

Nicely done! Looking forward to the pluggable storage engine.

pas8y ago

The JIRA tickets don't really shine with much hope :/

And the Rocksandra fork is already ~3500 commits behind master, so upstreaming this will be interesting.

Oh, and the Rocksandra fork is already kind of abandoned - no commits since 2017 Dec. (which probably means this is not actually the code that runs under Instagram.)

dikanggu8y ago

This is the rocksandra branch, https://github.com/Instagram/cassandra/tree/rocks_3.0, we develop it on top of Cassandra 3.0. It's the code we are running on our production servers.

1 more reply

jjirsa8y ago

I'm a committer, I'm familiar with the JIRA ticket.

1 more reply

kiril-me8y ago

It would be great. But I don't think it could happen. The pluggable storage engine would greatly increase the cognitive complexity of the code.

jjirsa8y ago

rbranson8y ago

dikanggu8y ago

Yeah, for Rocksandra, we are able to use much smaller heap size, and most of the objects are recycled during the young gen GC.

agnivade8y ago

> We also observed that the GC stalls on that cluster dropped from 2.5% to 0.3%, which was a 10X reduction!

Umm .. shouldn't the stalls go to 0, because now you have moved to C++ ? Or is this the time it takes for the manual garbage collection to occur ?

steeve8y ago

Why not use ScyllaDB ? (Serious)

cnlwsu8y ago

Answered a bit before but they have a team that knows c* well. Cassandra is proven to handle petabytes at scale in production systems.

xuanyue8y ago

Is there any trade off after replacing LSM tree-based storage engine to RocksDB storage engine?

irfansharif8y ago

RocksDB is also an LSM structured KV store.

welder8y ago

https://github.com/datastax/python-driver

ismail8y ago

So question:

Any thoughts on replacing HDFS + Yarn + Hive + HBASE with GulsterFS + Kubernetes + Cassandra

ddorian438y ago

Hbase is sync+globally sorted, while cassandra is not, so probably not.

alsadi8y ago

Can we add lz4 to the blend to reduce disk IO?

j / k navigate · click thread line to collapse