Astro - Hacker News

91 comments

nthypes an hour ago

https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main...
Model was released and it's amazing. Frontier level (better than Opus 4.6) at a fraction of the cost.
[-]
- 0xbadcafebee 18 minutes ago
  
  I don't think we need to compare models to Opus anymore. Opus users don't care about other models, as they're convinced Opus will be better forever. And non-Opus users don't want the expense, lock-in or limits.
  As a non-Opus user, I'll continue to use the cheapest fastest models that get my job done, which (for me anyway) is still MiniMax M2.5. I occasionally try a newer, more expensive model, and I get the same results. I have a feeling we might all be getting swindled by the whole AI industry with benchmarks that just make it look like everything's improving.
- NitpickLawyer 19 minutes ago
  
  > (better than Opus 4.6)
  There we go again :) It seems we have a release each day claiming that. What's weird is that even deepseek doesn't claim it's better than opus w/ thinking. No idea why you'd say that but anyway.
  Dsv3 was a good model. Not benchmaxxed at all, it was pretty stable where it was. Did well on tasks that were ood for benchmarks, even if it was behind SotA.
  This seems to be similar. Behind SotA, but not by much, and at a much lower price. The big one is being served (by ds themselves now, more providers will come and we'll see the median price) at 1.74$ in / 3.48$ out / 0.14$ cache. Really cheap for what it offers.
  The small one is at 0.14$ in / 0.28$ out / 0.028$ cache, which is pretty much "too cheap to matter". This will be what people can run realistically "at home", and should be a contender for things like haiku/gemini-flash, if it can deliver at those levels.
- onchainintel an hour ago
  
  How does it compare to Opus 4.7? I've been immersed in 4.7 all week participating in the Anthropic Opus 4.7 hackathon and it's pretty impressive even if it's ravenous from a token perspective compared to 4.6
  [-]
  - greenknight an hour ago
    
    The thing is, it doesnt need to beat 4.7. it just needs to do somewhat well against it.
    This is free... as in you can download it, run it on your systems and finetune it to be the way you want it to be.
    
    [-]
    
    p1esk an hour ago
    
    Do you think a lot of people have “systems” to run a 1.6T model?
    
    [-]
    
    applfanboysbgon 24 minutes ago
    
    No, but businesses do. Being able to run quality LLMs without your business, or business's private information, being held at the mercy of another corp has a lot of value.
    
    [-]
    
    forrestthewoods 4 minutes ago
    
    What type of system is needed to self host this? How much would it cost?
    
    choldstare 14 minutes ago
    
    Not really - on prem llm hosting is extremely labor and capital intensive
    
    [-]
    
    applfanboysbgon 9 minutes ago
    
    But can be, and is, done. I work for a bootstrapped startup that hosts a DeepSeek v3 retrain on our own GPUs. We are highly profitable. We're certainly not the only ones in the space, as I'm personally aware of several other startups hosting their own GLM or DeepSeek models.
    
    onchainintel 33 minutes ago
    
    Completely agree, not suggesting it needs ot just genuinely curious. Love that it can be run locally though. Open source LLMs punching back pretty hard against proprietary ones in the cloud lately in terms of performance.
    
    kelseyfrog 42 minutes ago
    
    What's the hardware cost to running it?
    
    [-]
    
    bbor 2 minutes ago
    
    I was curious, and some [intrepid soul](https://wavespeed.ai/blog/posts/deepseek-v4-gpu-vram-require...) did an analysis. Assuming you do everything perfectly and take full advantage of the model's MoE sparsity, it would take:
    - To run at full precision: "16–24 H100s", giving us ~$400-600k upfront, or $8-12/h from [us-east-1](https://intuitionlabs.ai/articles/h100-rental-prices-cloud-c...).
    - To run with "heavy quantization" (16 bits -> 8): "8xH100", giving us $200K upfront and $4/h.
    - To run truly "locally"--i.e. in a house instead of a data center--you'd need four 4090s, one of the most powerful consumer GPUs available. Even that would clock in around $15k for the cards alone and ~$0.22/h for the electricity (in the US).
    Truly an insane industry. This is a good reminder of why datacenter capex from since 2023 has eclipsed the Manhattan Project, the Apollo program, and the US interstate system combined...
    
    redox99 30 minutes ago
    
    Probably like 100 USD/hour
    
    slashdave 36 minutes ago
    
    "if you have to ask..."
    
    johnmaguire an hour ago
    
    ... if you have 800 GB of VRAM free.
    
    [-]
    
    inventor7777 29 minutes ago
    
    I remember reading about some new frameworks have been coming out to allow Macs to stream weights of huge models live from fast SSDs and produce quality output, albeit slowly. Apart from that...good luck finding that much available VRAM haha
  - rvz an hour ago
    
    It is more than good enough and has effectively caught up with Opus 4.6 and GPT 5.4 according to the benchmarks.
    It's about 2 months behind GPT 5.5 and Opus 4.7.
    As long as it is cheap to run for the hosting providers and it is frontier level, it is a very competitive model and impressive against the others. I give it 2 years maximum for consumer hardware to run models that are 500B - 800B quantized on their machines.
    It should be obvious now why Anthropic really doesn't want you to run local models on your machine.
    
    [-]
    
    deaux 6 minutes ago
    
    Vibes > Benchmarks. And it's all so task-specific. Gemini 3 has scored very well in benchmarks for very long but is poor at agentic usecases. A lot of people prefering Opus 4.6 to 4.7 for coding despite benchmarks, much more than I've seen before (4.5->4.6, 4->4.5).
    Doesn't mean Deepseek v4 isn't great, just benchmarks alone aren't enough to tell.
    
    snovv_crash 13 minutes ago
    
    With the ability of the Qwen3.6 27B, I think in 2 years consumers will be running models of this capability on current hardware.
    
    colordrops 26 minutes ago
    
    What's going to change in 2 years that would allow users to run 500B-800B parameter models on consumer hardware?
    
    [-]
    
    DiscourseFan 14 minutes ago
    
    I think its just an estimate
- doctoboggan an hour ago
  
  Is it honestly better than Opus 4.6 or just benchmaxxed? Have you done any coding with an agent harness using it?
  If its coding abilities are better than Claude Code with Opus 4.6 then I will definitely be switching to this model.
  [-]
  - madagang 40 minutes ago
    
    Their Chinese announcement says that, based on internal employee testing, it is not as good as Opus 4.6 Thinking, but is slightly better than Opus 4.6 without Thinking enabled.
    
    [-]
    
    deaux 5 minutes ago
    
    That's super interesting, isn't Deepseek in China banned from using Anthropic models? Yet here they're comparing it in terms of internal employee testing.
    
    mchusma 34 minutes ago
    
    I appreciate this, makes me trust it more than benchmarks.
- bbor 17 minutes ago
  
  For the curious, I did some napkin math on their posted benchmarks and it racks up 20.1 percentage point difference across the 20 metrics where both were scored, for an average improvement of about 2% (non-pp). I really can't decide if that's mind blowing or boring?
  Claude4.6 was almost 10pp better at at answering questions from long contexts ("corpuses" in CorpusQA and "multiround conversations" in MRCR), while DSv4 was a staggering 14pp better at one math challenge (IMOAnswerBench) and 12pp better at basic Q&A (SimpleQA-Verified).
  [-]
  - Quasimarion 13 minutes ago
    
    FWIW it's also like 10x cheaper.
- sergiotapia an hour ago
  
  The dragon awakes yet again!
  [-]
  - kindkang2024 16 minutes ago
    
    There appears a flight of dragons without heads. Good fortune.
    That's literally what the I Ching calls "good fortune."
    Competition, when no single dragon monopolizes the sky, brings fortune for all.
- rapind an hour ago
  
  Pop?
simonw 22 minutes ago

I like the pelican I got out of deepseek-v4-flash more than the one I got from deepseek-v4-pro.
Flash: https://gist.github.com/simonw/4a7a9e75b666a58a0cf81495acddf...
Pro: https://gist.github.com/simonw/9e8dfed68933ab752c9cf27a03250...
Both generated using OpenRouter.
[-]
- nickvec 16 minutes ago
  
  The Flash one is pretty impressive. Might be my favorite so far in the pelican-riding-a-bicycle series
- JSR_FDED 19 minutes ago
  
  No way. The Pro pelican is fatter, has a customized front fork, and the sun is shining! He’s definitely living the best life.
  [-]
  - w4yai 15 minutes ago
    
    yeah. look at these 4 feathers (?) on his bum too.
- ycui1986 10 minutes ago
  
  I really like the pro version. The pelican is so cute.
yanis_t 22 minutes ago

Already on Openrouter. Pro version is $1.74/m/input, $3.48/m/output, while flash $0.14/m/input, 0.28/m/output.
[-]
- astrod 2 minutes ago
  
  Getting 'Api Error' here :( Every other model is working fine.
- esafak 11 minutes ago
  
  https://openrouter.ai/deepseek/deepseek-v4-pro
  https://openrouter.ai/deepseek/deepseek-v4-flash
  [-]
  - 77ko 2 minutes ago
    
    Its on OR - but currently not available on their anthropic endpoint. OR if you read this, pls enable it there! I am using kimi-2.6 with Claude Code, works well, but Deepseek V4 gives an error:
    `https://openrouter.ai/api/messages with model=deepseek/deepseek-v4-pro, OR returns an error because their Anthropic-compat translator doesn't cover V4 yet. The Claude CLI dutifully surfaces that error as "model...does not exist"
seanobannon an hour ago

Weights available here: https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro
mchusma 14 minutes ago

For comparison on openrouter DeepSeek v4 Flash is slightly cheaper than Gemma 4 31b, more expensive than Gemma 4 26b, but it does support prompt caching, which means for some applications it will be the cheapest. Excited to see how it compares with Gemma 4.
fblp an hour ago

There's something heartwarming about the developer docs being released before the flashy press release.
[-]
- necovek 36 minutes ago
  
  Where's the training data and training scripts since you are calling this open source?
- onchainintel an hour ago
  
  Insert obligatory "this is the way" Mando scene. Indeed!
zargon an hour ago

The Flash version is 284B A13B in mixed FP8 / FP4 and the full native precision weights total approximately 178 GB. KV cache is said to take 10% as much space as V3. This looks very accessible for people running "large" local models. It's a nice follow up to the Gemma 4 and Qwen3.5 small local models.
[-]
- sbinnee 25 minutes ago
  
  Price is appealing to me. I have been using gemini 3 flash mainly for chat. I may give it a try.
  input: $0.14/$0.28 (whereas gemini $0.5/$3)
  Does anyone know why output prices have such a big gap?
jessepcc an hour ago

At this point 'frontier model release' is a monthly cadence, Kimi 2.6 Claude 4.6 GPT 5.5 — the interesting question is which evals will still be meaningful in 6 months.
sidcool 35 minutes ago

Truly open source coming from China. This is heartwarming. I know if the potential ulterior motives.
[-]
- try-working 5 minutes ago
  
  if you want to understand why labs open source their models: http://try.works/why-chinese-ai-labs-went-open-and-will-rema...
- I_am_tiberius 18 minutes ago
  
  Open weight!
jdeng an hour ago

Excited that the long awaited v4 is finally out. But feel sad that it's not multimodal native.
gbnwl an hour ago

I’m deeply interested and invested in the field but I could really use a support group for people burnt out from trying to keep up with everything. I feel like we’ve already long since passed the point where we need AI to help us keep up with advancements in AI.
[-]
- wordpad an hour ago
  
  The players barely ever change. People don't have problems following sports, you shouldn't struggle so much with this once you accept top spot changes.
  [-]
  - ehnto 16 minutes ago
    
    It is funny seeing people ping pong between Anthropic and ChatGPT, with similar rhetoric in both directions.
    At this point I would just pick the one who's "ethics" and user experience you prefer. The difference in performance between these releases has had no impact on the meaningful work one can do with them, unless perhaps they are on the fringes in some domain.
    Personally I am trying out the open models cloud hosted, since I am not interested in being rug pulled by the big two providers. They have come a long way, and for all the work I actually trust to an LLM they seem to be sufficient.
    
    [-]
    
    DiscourseFan 11 minutes ago
    
    I find ChatGPT annoying mostly
    
    [-]
    
    awakeasleep 7 minutes ago
    
    Open settings > personalization. Set it to efficient base style. Turn off enthusiasm and warmth. You’re welcome
mariopt 8 minutes ago

Does deepseek has any coding plan?
aliljet 22 minutes ago

How can you reasonably try to get near frontier (even at all tps) on hardware you own? Maybe under 5k in cost?
[-]
- awakeasleep 9 minutes ago
  
  The same way you fit a bucket wheel excavator in your garage
- jdoe1337halo 7 minutes ago
  
  More like 500k
taosx an hour ago

MErge? https://news.ycombinator.com/item?id=47885014
luyu_wu an hour ago

For those who didn't check the page yet, it just links to the API docs being updated with the upcoming models, not the actual model release.
[-]
- talim an hour ago
  
  Weights are on Huggingface FWIW. https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/tree/main
- cmrdporcupine an hour ago
  
  My submission here https://news.ycombinator.com/item?id=47885014 done at the same time was to the weights.
  dang, probably the two should be merged and that be the link
  [-]
  - culi 37 minutes ago
    
    there's no pinging. Someone's gotta email dang
Aliabid94 an hour ago

MMLU-Pro:
Gemini-3.1-Pro at 91.0
Opus-4.6 at 89.1
GPT-5.4, Kimi2.6, and DS-V4-Pro tied at 87.5
Pretty impressive
[-]
- ant6n 4 minutes ago
  
  Funny how Gemini is theoretically the best -- but in practice all the bugs in the interface mean I don't want to use it anymore. The worst is it forgets context (and lies about it), but it's very unreliable at reading pdfs (and lies about it). There's also no branch, so once the context is lost/polluted, you have to start projects over and build up the context from scratch again.
namegulf 29 minutes ago

Is there a Quantized version of this?
KaoruAoiShiho an hour ago

SOTA MRCR (or would've been a few hours earlier... beaten by 5.5), I've long thought of this as the most important non-agentic benchmark, so this is especially impressive. Beats Opus 4.7 here
swrrt an hour ago

Any visualised benchmark/scoreboard for comparison between latest models? DeepSeek v4 and GPT-5.5 seems to be ground breaking.
rvz an hour ago

The paper is here: [0]
Was expecting that the release would be this month [1], since everyone forgot about it and not reading the papers they were releasing and 7 days later here we have it.
One of the key points of this model to look at is the optimization that DeepSeek made with the residual design of the neural network architecture of the LLM, which is manifold-constrained hyper-connections (mHC) which is from this paper [2], which makes this possible to efficiently train it, especially with its hybrid attention mechanism designed for this.
There was not that much discussion around it some months ago here [3] about it but again this is a recommended read of the paper.
I wouldn't trust the benchmarks directly, but would wait for others to try it for themselves to see if it matches the performance of frontier models.
Either way, this is why Anthropic wants to ban open weight models and I cannot wait for the quantized versions to release momentarily.
[0] https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main...
[1] https://news.ycombinator.com/item?id=47793880
[2] https://arxiv.org/abs/2512.24880
[3] https://news.ycombinator.com/item?id=46452172
[-]
- jeswin an hour ago
  
  > this is why Anthropic wants to ban open weight models
  Do you have a source?
reenorap 38 minutes ago

Which version fits in a Mac Studio M3 Ultra 512 GB?
[-]
- simonw 20 minutes ago
  
  The Flash one should - it's 160GB on Hugging Face: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash/tree/ma...
  [-]
  - ycui1986 6 minutes ago
    
    So, dual RTX PRO 6000
ls612 an hour ago

How long does it usually take for folks to make smaller distills of these models? I really want to see how this will do when brought down to a size that will run on a Macbook.
[-]
- simonw 18 minutes ago
  
  Unsloth often turn them around within a few hours, they might have gone to bed already though!
  Keep an eye on https://huggingface.co/unsloth/models
  Update ten minutes later: https://huggingface.co/unsloth/DeepSeek-V4-Pro just appeared but doesn't have files in yet, so they are clearly awake and pushing updates.
- inventor7777 25 minutes ago
  
  Weren't there some frameworks recently released to allow Macs to stream weights from fast SSDs and thus fit way more parameters than what would normally fit in RAM?
  I have never tried one yet but I am considering trying that for a medium sized model.
  [-]
  - simonw 10 minutes ago
    
    I've been calling that the "streaming experts" trick, the key idea is to take advantage of Mixture of Expert models where only a subset of the weights are used for each round of calculations, then load those weights from SSD into RAM for each round.
    As I understand it if DeepSeek v4 Pro is a 1.6T, 49B active that means you'd need just 49B in memory, so ~100GB at 16 bit or ~50GB at 8bit quantized.
    v4 Flash is 284B, 13B active so might even fit in <32GB.
    
    [-]
    
    inventor7777 7 minutes ago
    
    Ahh, that actually makes more sense now. (As you can tell, I just skimmed through the READMEs and starred "for later".)
    My Mac can fit almost 70B (Q3_K_M) in memory at once, so I really need to try this out soon at maybe Q5-ish.
  - the_sleaze_ 21 minutes ago
    
    Do you have the links for those? Very interested
    
    [-]
    
    inventor7777 17 minutes ago
    
    Sure!
    Note: these were just two that I starred when I saw them posted here. I have not looked seriously at it at the moment,
    https://github.com/danveloper/flash-moe
    https://github.com/t8/hypura
hongbo_zhang 22 minutes ago

congrats
frozenseven an hour ago

Better link:
https://news.ycombinator.com/item?id=47885014
https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro
nickandbro an hour ago

Very impressive throughput performance
shafiemoji an hour ago

I hope the update is an improvement. Losing 3.2 would be a real loss, it's excellent.
raincole an hour ago

History doesn't always repeat itself.
But if it does, then in the following week we'll see DeepSeek4 floods every AI-related online space. Thousands of posts swearing how it's better than the latest models OpenAI/Anthropic/Google have but only costs pennies.
Then a few weeks later it'll be forgotten by most.
[-]
- sbysb 37 minutes ago
  
  It's difficult because even if the underlying model is very good, not having a pre-built harness like Claude Code makes it very un-sticky for most devs. Even at equal quality, the friction (or at least perceived friction) is higher than the mainstream models.
  [-]
  - raincole 32 minutes ago
    
    OpenCode? Pi?
    If one finds it difficult to set up OpenCode to use whatever providers they want, I won't call them 'dev'.
    The only real friction (if the model is actually as good as SOTA) is to convince your employer to pay for it. But again if it really provides the same value at a fraction of the cost, it'll eventually cease to be an issue.
  - cmrdporcupine 29 minutes ago
    
    They have instructions right on their page on how to use claude code with it.