Upcoming 0.18 upgrade, 404 errors and infrastructure costs

sunaurus@lemm.ee · edit-2 1 year ago

Upcoming 0.18 upgrade, 404 errors and infrastructure costs

two_wheel2@lemm.ee · edit-2 1 year ago

Alright, I’m tossing my tiny hat into the sponsor ring. Thanks so much for putting this community together! I’m excited to see it grow. Just out of curiosity, what does the incremental cost look like? Does it scale well with users? Or does it explode a little bit?

sunaurus@lemm.ee · edit-2 1 year ago

That’s greatly appreciated!

In terms of costs of scaling, I would say we’re positioned a bit better than many other Lemmy instances at the moment, thanks to the fact that we employ horizontal scaling as much as possible for the Lemmy software itself.

By the way, AFAIK, lemm.ee is the only non-experimental Lemmy instance that has chosen to go with horizontal scaling so far. If anybody knows of any other instance that is doing it, I would be super interested to know about it! All the admins I’ve spoken to so far myself have confirmed that they are only doing vertical scaling.

More technical details below for anybody who is interested:

There are two approaches you can generally take for scaling - horizontal, where you add more load balanced nodes of more or less the same power, or vertical, where you increase the power of an individual node (of course a mixture of both is also possible).

One of the benefits of horizontal scaling is that in most cases, it’s significantly more flexible compared to vertical scaling. For example, at my current cloud provider, the only upgrade path for vertical scaling a server would be 8 CPU -> 16 CPU - 32 CPU -> 40 CPU. So if you’re on a 16 CPU server, and you need just a little bit more headroom, then your only option is to upgrade to the 32CPU server, which is straight up double the power (and cost!). Meanwhile, with horizontal scaling, you can just keep adding smaller servers (say 2 CPU each) one at a time, thus growing costs more gradually and appropriately for your actual needs.

So for lemm.ee, this horizontal scaling means that when our backend servers start getting overloaded, I can just add one or two more servers without exponentially increasing costs.

OneDimensionPrinter@lemm.ee · 1 year ago

As someone who has “been there and done that” at a much larger scale than many devs may ever get a chance to (not a brag, it can suck royally) this really seems like the smart choice.

This is effectively a basic web server scenario and horizonal scaling tends to with really well to a point. And frankly it’ll be a long while before that becomes the bottleneck.

Smart choices you’re making. All the best and I’m happy to help out monetarily where I can!

electromage@lemm.ee · 1 year ago

Are your servers in one geographic region? Could you scale across regions for better performance?

sunaurus@lemm.ee · 1 year ago

I am already leveraging Cloudflare’s globally distributed cache, which helps improve performance even if you’re far away from the backend server. But this only helps partially, not with all types of requests.

lemm.ee is hosted in central Europe, and based on monitoring, it does seem that most users are having a pretty decent experience on lemm.ee regardless of their geographic location so far. One key exception to this are short windows of database load spikes, which last for roughly 10 seconds every 5 minutes. For these spikes, everybody is suffering equally, regardless of where they are in the world 😅.

But in general I agree with the sibling comment by @Notorious - rather than scaling one instance to be some massive globally distributed powerhouse, it makes sense to spread out the load amongst a lot of different instances.

OneDimensionPrinter@lemm.ee · 1 year ago

Are the DB spikes ACTUALLY every 5 minutes or is that just kind of a guess? I ask because if it’s consistent, it’s gotta be some sidecar process somewhere in the stack that can be fiddled with.

That said, it really sounds like you know what you’re doing already so I’ll just go play with my new communities.

sunaurus@lemm.ee · 1 year ago

The spikes are caused by a specific reoccurring process which happens every 5 minutes. I have already significantly optimized it with a patch on lemm.ee, I’m working on getting it merged upstream as well!

xavier666@lemm.ee · 1 year ago

For storage, I can understand how horizontal scaling works (add more storage nodes to, say glusterfs). But how does it work for CPU? Since adding a 2CPU VM can be physically on another server, it would need lemmy to work in a highly distributed manner, i.e., CPU instructions need to cross the network.

Is this distributed feature a part of lemmy or is there another abstraction layer?

sunaurus@lemm.ee · edit-2 1 year ago

This is where our load balancer comes in. All requests go through the load balancer, and this load balancer will try to evenly distribute the requests to all of our backend servers.

Is this distributed feature a part of lemmy … ?

In fact it’s the opposite - Lemmy has so far had some assumptions built in to the code which make it quite hard to run on multiple servers. I have made some modifications in order to improve this (and contributed those modifications back to the main repo as well). It’s one of the things I want to keep improving as we grow.

bric@lemm.ee · 1 year ago

Same. I’m not putting in a ton, but monthly donations go a long way to help with monthly server costs. We just need 150 people to put in $1 a month and we’ll be covered indefinitely

two_wheel2@lemm.ee · 1 year ago

Exactly. I’ve tossed in $5/mo and I literally just realized that with Reddit never in my WILDEST DREAMS would I have imagined kicking in some money for something like gold or trophies or even Apollo (RIP), but $5 a month to contribute to supporting a distributed community of people beyond myself feels like nothing to me. I think that speaks to the potential federation + good will can offer the world

Upcoming 0.18 upgrade, 404 errors and infrastructure costs

Upcoming 0.18 upgrade, 404 errors and infrastructure costs

Hello, fellow lemmings!

Upcoming 0.18 upgrade

Why do we even want 0.18?

Random 404 errors

Server costs

Pinning updates on the front page