Yeah it was wild watching how many contractors came out when Google was rolling their fiber through my neighborhood, like it was different contractor groups for every step of the way. I'm not complaining a 5gbit connection is quite nice.
I'm incredibly lucky (even in the US) that I have symmetric 5gbit connection and can go up to 8gbit if I wanted (but I don't have everything wired for 10gbit so it doesn't make sense right now). I'd love to see more of that in the US, but know it's not happening anytime soon.
I have 5Gbps symmetrical. I love having it, but the biggest use-case has been bragging about it to my colleagues. I haven't even had the full bandwidth available at my desk for a while now (I've had a terrible time terminating cat6a for reliable 10Gbe, so I'm usually on wifi), and it hasn't really bothered me much.
For me, with upgraded networking it's been huge. Much faster downloads and uploads, I can share more with family members, and when multiple people are using my bandwidth, no slowdowns. Most of my equipment is 2.5g, but still is a fairly significant upgrade, and the per month additional cost is $30/month so makes it worth it to me.
I have had an impossible time getting 120B or better models running on Strix Halo (especially under Windows) with any large context windows. And 30-40 tokens/second is fine, but not the fastest.
For the most part lately I have been sticking with Qwen 3.8 27b and that thing will easily suck up 64gb of ram. Add in docker with some additional programs running and it's really easy to eat up 128gb of ram.
I found a couple of different 4-bit quantizations of Laguna S 2.1 that run pretty well with pretty big context (also quantized, to 8 bits, I think). Unfortunately, Laguna isn't better than Qwen 3.8 27B, which I'm able to run at roughly the same speed on my desktop machine, so I don't use Laguna or the Strix Halo very much, lately. (It's also too hot for me to be running heaters for inference. It's been ~110F most days for the past few weeks.)
Aren't you kind of stuck with smaller models on the mini though? Even with the pro you'd be stuck with a max of four daisy chained over thunderbolt and with the 160 gig memory bandwidth you'd probably be far better off with other configurations.
I really haven't seen any 16gb Mac Mini cluster setups running large models at any appreciable speed for the reasons previously provided. Do you happen to have examples that folks are actually using?
A Pi isn't a NAS with an old CPU and minimal RAM which is running a bunch of other things as well. I have run Plex on my Synology, and it's fine, but I also have 8gb ram and honestly always have a better experience running Jellyfin on a small PC. Easier that way, and you can run other services. Before the rampocalypse you could find mini-PCs with 32gb ram and 1tb SSD for like $350, and have great transcoding ability when you are on the go or have multiple streams going.
As long as everything is using 10gb, faster transfer speeds when moving data between computers. For downloads, it wont help unless you have 10gb fiber, but for most folks, 2.5gb is quite fast. Hell my spinning media NAS has a hard time saturating when moving files between internal servers.
I have had a difficult time with running 120b models on my 128gb setup, especially with any larger context size. The 6bit of Qwen 3.5 is already just over 100gb, and when you go down to 4bit it seems a bit lobotomized.
For me it's be Strix Halo, 128gb machine, especially running Qwen models. Except when I bought it, it was $1,900, now it's $4,600 for the same box. (Wow that's insane)
For tinkering and learning, it's been great. Tie it into something like Hermes and you have a pretty powerful AI assistant in a box. And when you need to step up your model, you just do something like OpenRouter and it makes it pretty easy.