Hacker Newsnew | past | comments | ask | show | jobs | submit | karmakaze's commentslogin

Personally I'm using Qwen3.8-27B (MXFP4 quant W4A8) locally hosted on a pair of AMD GPUs (with DeepSeek Harness). It starts at 250 tokens/sec down to 120 past 128k context.

At work mostly Opus 4.8 (sometimes a GPT or Gemini 3.1 Pro). I find Opus 5 chatty/slower and Fable can venture into over-engineering itself into unnecessary complications.


Busabase[0]

> Languages: TypeScript 98.4%, Other 1.6%

Not for me.

[0] https://github.com/busabase/busabase


What lang you want?

Seems like a random rant to me.

> What is the internet now to me? In many ways it’s something that I have to tolerate to do many things that functioned fine before. I have to use an app to pay for parking. I need to create an online account to pay for a family swim at the local leisure centre. I have to submit an online form to book a slot at the local recycling centre. I need an app to check my bank and credit card balances.

It's so much more annoying to do any of these things without/before internet. I don't do half of the other things in the rest of that paragraph. Like for phones I just get a used CAD$400 Android phone with good audio output.


> Seems like a random rant to me

I think this applies to all personal blog posts


I think it's largely due to psychology. If 210m is considered the best, then as they approach it they may start to tense up and choke. When the goal and possibility is known as 500m, then there's no point being concerned near 210m.

The way it does tensor splitting without all-reduce cost over PCIe bus wasn't something I thought was possible.

What kind of performance are you getting with 4x R9700s--what do you do with all the VRAM (batching, concurrent requests, etc)?


personally? i have 2x gpus.. but i get bursts of ~200tok/s generation, and around 4500-5000tok/s prefil

Yeah the R4D Kernel rules imho.


Similar here peak ~250 and down to ~120 as it gets close to 128k (which is where I set DSH compaction) though it can readily do 256k.

I just got DeepSeek Harness (DSH) set up with 2x R9700 and it's rather mind blowing that these can do actual work and quickly. Up until now I've always been evaluating and searching for better hardware/model/tweaks. This is much more than I even hoped for and considered getting extra 3090/4090. Now I can stop looking/tweaking and start using it for all the different things I've yet to discover it's good for. I do plan to also try/use Hermes and Pi. DSH is annoying that every plugin install/remove requires a restart--given that "everything's a plugin".


Thanks! Didn't expect to see this here. Exactly what I needed to run Qwen3.8-27B-Quark-AWQ-MXFP4-native.gguf as well as other experiments on one or 2x R9700's (I hope).

> Anything design-y was “too Apple” and unacceptably bourgeois.

That was definitely the case of the Unity desktop that really only worked well on netbooks. That's when I lost confidence in Ubuntu for design. Loss in Canonical on the whole came later.


The previous 'personal' AI Station I had pictured was the a16z one[0].

Each MI350P[1] in the TR Halo Station has 144GB VRAM and with 4.6 PFLOPs peak MXFP6 performance.

Four liquid cooled? Yes please.

[0] https://a16z.com/building-a16zs-personal-ai-workstation-with...

[1] https://www.amd.com/en/products/accelerators/instinct/mi350/...


NasaFTs are back in style!

It seems we could use a new kind of memory that streams the weight data in, like GDDR in reverse.


Optane? (Too soon?)

Optane was targeting the latency gap, while HBF targets bandwidth. Given how LLMs work, HBF is perfect for offloading.

interesting!

yes! I guess future hardware designs will have something like that!

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: