Insane scores for a model of this size. But it does seem to be a rather insane over-thinker with the biggest token use of any model, which combined with 256k context size it means it wont do much before filling it context
We don't actually know how much thinking GPT and Opus do, the labs won't show us anymore. And they certainly take their time before starting to answer.
completely vibes based, but ive been using it to port Mindustry game from Java to C# with agents, and its been working for 50 hours (its 15-20 tks so super slow inference). Its done a fantastic work and its almost finished now. Better results than deepseek flash and gpt luna by a mile on this kind of long term work. Less good than gpt sol or opus. We dont know the param count but my guess is 200-300 range.
FWIW on a 6000 Pro it takes 68 seconds for 10 seconds 480p video, (cold) same demo workflow as you used. Set "megapixels" to 2.0 (1920x1088) and same video seems to take 5+ minutes, not sure if everything is right/correct at the moment.
As far as I can tell, the current ComfyUI nodes don't even do compilation, and I haven't looked into what attention mechanism they're using, but I'm sure with time these durations will come down even more.
People on reddit have definitely pointed out that sageattention will speed up the renders.
And it's literally the first day. Someone will make a distilled 4-8 step LoRA and we're off to the races.
Edit: did a couple of 10 second long 864/480 i2v videos on my RTX Pro 6000: sageattention bumps them up 33%, that is to say, 140.89 seconds without sageattention becomes 105.69 with sageattention on (if using the KJ Sageattention node, "allow_compile" doesn't seem to affect it, just "sage_attention" set to "auto" works fine).
EasyCache also appears to work, but does affect quality, at least with the default threshold or even down to 0.10. Still, at 0.10 threshold the same render above, with sageattention, is down to 71.33 seconds, so depending on your use case the quality hit might be worth it. Also it seems that with EasyCache the video still matches the un-EasyCached video (with the same seed), so you could use it to do seed hunting.
I'm on an RTX 5090. I told it to make a 5 second 864x480 video, it's been running for over 30 minutes and is only 35% done in the SamplerCustomAdvanced step.
EDIT: Oh, I'm an idiot. Forgot I had a llama.cpp webserver running with a model loaded. Killed it and it finished very fast.
I am particularly curious how multimodal models will work with types of knowledge that are inherently non-text. For example, SOTA LLMs really suck at electronics, especially analog electronics.
Is MiniMax H3 capable of logical / technical reasoning, or is it purely art oriented?
Suck at what aspect of (analog) electronics specifically? Not contradicting the claim, just want to understand it.
I have not tested yet, but I suspect that LLMs with a harness that can execute code can do SPICE simulations rather ok these days? I have seen MCPs for measurement equipment also, maybe they can even close the physical loop?
> I have seen MCPs for measurement equipment also, maybe they can even close the physical loop?
I'm working on this but for various reasons can't have my physical lab up at the moment. But yes there are many options for connected test equipment that could rather trivially be interacted with via LLM or pretty easy to write libraries.
spice is a very bad simulation. it's not like a unit test or a VM. it works very differently to the real thing, only approximating it in some highly controlled situations.
If you have accurate device models, and you know how to use simulations, you can predict quite well the behavior of the analog circuits that you design.
Simulations that work like the real thing are an absolute necessity in the industry. It is frequent to design the analog parts of some integrated circuit, when you might have to wait months to get prototype samples that you can measure in the laboratory.
If the prototypes do not work exactly as expected, you will probably have the chance to make once a correction of the schematics and/or layout, which will be incorporated in the final product.
But when not even the second try does not work as designed, that is likely to be a failure, because any other redesign would add months and many millions of $ until the product can reach the market.
So "spice" simulations that can be trusted are a necessity. When designing integrated circuits you cannot use empirical methods that work when you make manually a single-use board, where you can replace components or adjust resistors, capacitors or inductors until you get what you want. For integrated circuits, you must know in advance how the circuit will work, because you cannot tweak it.
It is true however that it is easy to configure simulations that will give non-sense results, due to lack of convergence or other numerical errors, so to avoid this, one typically needs both experience with laboratory measurements and experience in using the simulator.
I doubt that an LLM would successfully run SPICE simulations, because there are few examples of written instructions of how to do this in the right way (which varies depending on the domain of applications, and each SPICE-like simulator may have its own quirks).
> work with types of knowledge that are inherently non-text
What are those things exactly? AFAIK, everything we can "know" can be written down, one way or another, even analog circuits.
Also, what SOTA LLMs are you referring to? GPTs been handling analog circuits fine for quite some time, I want to say for at least one year? I've been "pair programming" a bunch of working circuits with GPT models since probably GPT 5 or so.
> AFAIK, everything we can "know" can be written down, one way or another, even analog circuits.
writing down circuit diagrams is like cooking about music.
> what SOTA LLMs are you referring to?
I have done a survey among analog electronics designers just a couple weeks ago and they all said that their forays into LLMs were great for digital electronics, code, and firmware, but for analog they were pretty terrible, with a variety of LLMs, according to everyone.
I think that's different though, that's "doing" rather than just "knowing". You can ride a bike without knowing how it works, and obviously vice-versa too. I don't see circuit diagrams as "doing" though, but the soldering part of building circuits definitely is that way though, you can't just read about it and excel first time you pick up an iron, you have to practice and understand it with your body, like bicycling.
Im the author of vkguide.dev and i disagree. I wrote that tutorial specifically for people that have already gone through learnopengl and want to learn vulkan.
The core problem with the modern apis is that they are so incredibly complicated that they will kill any newbie on the spot. When the student still doesnt really know what a mesh is, having them setup GPU side memory allocators and graphics compute pipelines is absurd.
Opengl meanwhile is significantly easier, and will give them the important terminology and math required to do graphics. Once the student has learned opengl and wrote a small renderer with a few basic techniques, moving to vulkan or DX12 will happen far more smoothly.
In particular, learnopengl explains a lot of basics like transformation coordinates, what a mesh is, and other similar "basic knowledge", while all vulkan and dx12 tutorials skip through that because they are meant for a much more experienced audience.
The problem with most modern "how to learn graphics programming" guides is that they over index on "low level GPU programming" and under index on important foundational 3D graphics concepts. Transformations and coordinate systems, model space, world space, view space, clip space, normalized device coordinates, screen space, concepts like parallel vs. perspective projection, viewport and viewport transformation, colors, lighting, materials, reflection, blending, antialiasing, texture mapping, various buffer types, tessellation...
You want to start with OpenGL because the API is organized roughly around these high level concepts. Starting with something like Vulkan means you're dropped straight into hundreds of lines of low level boilerplate setting up VkCommandBuffers and VkRenderPasses and VkPipelines and shaders, and all this GPU programming, and it's all just bewildering if you're used to thinking about things at the high level first, or just want to draw a fucking triangle on the screen.
Webgpu is the worst of both worlds. Significantly more complicated than opengl yet not capable of many "essential" modern techniques like bindless resources.
If you must use a modern api to learn for whatever reason, go with vulkan and use dynamic rendering and buffer device addressing.
WebGPU has the issue that feature/capability-wise it is essentially 5 years behind OpenGL 4.6, which came out 8 years ago. And the other downside being that it adopted old Vulkan concepts that even Vulkan started to ditch, like render passes and static pipelines.
For beginners that's irrelevant. WebGPU has more than enough features for beginners to learn about graphics programming. Not just that, once you're used to WebGPU, going to Vulkan will be much easier than the transition from OpenGL.
I wouldnt recommend anyone going to Vulkan, though. It's pretty much the worst graphics API out there, and WebGPU is mimicking outdated Vulkan design decisions that even Vulkan is currently outphasing, like render passes and static pipelines.
Ive been pleasantly surprised with the quality of this model. Its really good at code review and PR review. But right now due to them being so overloaded + this being a big model, its SO slow. It takes forever to do a relatively simple code review.
Oh, okay. Didn't know that, as I haven't tried K3 myself -- I'm quite content with Kimi K2.7-Code for the moment, which OpenRouter serves me to via US-based inference providers. Supposedly.
Obviously, not relevant if your reason for being here is to ferret out the excitement about K3 specifically. But if this is your first time trying Kimi, try K2.6 or K2.7! If my experience is any indication, you'll be blown away by those, too.
Deepseek V4 pro is a heavily undertrained model, they only trained it a bit more than the small version and that small version is 6-ish times smaller. Ive found that Flash is absolutely incredible as a workhorse for wide scale agentic nonsense, but Pro is a bit undercooked and really goes on strange tangents very often.
Yeah, I would have expected Zhipu to ship a Fable-adjacent model by the end of the year, but the jump from Kimi 2.7 (which I think is just barely at the level where it is genuinely helpful for coding) to this is absolutely bonkers. And this is clearly not just benchmaxing; this thing actually works.
If you told me I could only use this and never use Fable or Sol again, I'd shrug and not feel like I'd lost much.
Ive done this sort of thing using deepseek flash swarms for just a few bucks. The key thing is that each project needs its own swarm script. you cant just tell the bot "pls port this to rust". you need to analyze the architecture, build a proper plan of how the translation will go, and then iterate your execution script to launch hundreds of bots to do the various stages of the translation. And you need to iterate these stages to make sure it all works.
So, for example, i ported stb_image from C to Jai (you can think of Jai as similar to Zig, another c-style lang).
To do that, i built my own agent scheduling/swarm system, because there is nothing out there that works, its all slop.
First i used normal claudecode and the likes to build a set of translation rules, documentation about the project, and set up the context for all.
Then the first stage of the swarm was to use claude to split the 7000 line library into chunks of about 500 lines each, a few functions/structs at a time. Then each deepseek bot had a task of translating that snippet of code from C to Jai.
Next, i had a system merge it all back together. This was done with just a single claudecode running over time. It had to be agentic and not a script because the bots often would duplicate definitions and things like that.
Then, i had a system where it would try to compile the project, get an error list, parse the error list, and split that error list across multiple fixer agents.
Doing multiple fixers because Jai can give you a large error list across a bunch of functions, so this speeds it up. If your project rules/setup/etc is better you can improve this step to have less errors to begin with. The more tools you can have to prevent errors instantly on a feedback loop of each sub-task, the better it goes.
After the fixer step, the code was ported, but with some mistakes. I ran a wide pass of having 2 agents look at each struct/function, and validate if the translation was done well. 2 of them because if both agree the function is fine, the function probably is actually fine.
Next, i began iterating the test suite. The library had a decent enough test suite, so splitting the work between testing + fixing tests across the different parts of the library was doable. for example i would run fixes on png format and jpeg format at once couse their code doesnt overlap.
After the tests were 100% green, the next stage was to implement a literal exploit finding swarm to fuzz it all, which actually found a couple errors on the original library and found a bunch of more obscure bugs in the port.
After that, now i have my own version of this image loading library, more hardened than the original, on a pet lenguage i wanted to do it for.
Note that the better the AI you have, the less error the process has, and the larger chunking you can use. Mythos can likely do this same port much easier, but deepseek flash did it in about 3 bucks of tokens including the trial runs and experiments.
I later use this tooling to build a translator agent swarm for my videogame that can do english to anything translation for 7000 strings that can be menu buttons, tooltips, small fluff text, and others.
Every developer on videogames has some kind of offline mode already implemented, because its necessary to be able to playtest the game builds on the developer machines. Any argument against SKG is lobbyst nonsense. With the very specific exception of stuff like MMOs. We are seeing cases of pirates being able to play those "turned off games" through cracks and private servers, so there is absolutely 0 reason why the publisher cant already do it.
> Every developer on videogames has some kind of offline mode already implemented, because its necessary to be able to playtest the game builds on the developer machines.
Not guaranteed. Many just run a local server, either in-process or externally. Minecraft's singleplayer mode actually runs a server in-process internally. This simplifies development because singleplayer is conceptually the same as playing alone on a server.
This gets more complicated when there are infrastructure servers in the mix for things like player state, matchmaking, etc. You would bypass that in development but they are required for normal play while being external to the game server.