Meta Muse Glimmer – open weights 30B local coding model

(research.meta.ai)

567 points | by riordan 5 hours ago

45 comments

  • avaer 3 hours ago
    I lament the comments saying this in any way redeems Meta (the company).

    The researchers releasing this stuff have almost nothing to do with Meta other than being bankrolled by the slaughterhouse.

    You aren't the customer, you are the pawn in big tech's game of thrones. Your good will is a commodity to be traded, almost literally. It will be used against you the moment it's convenient. This is open weights because Meta couldn't monetize it in any other way than to cloud developer's judgement of their reputation.

    But I guess most people just don't care.

    I'm glad it's open. It does not make me think any better of Meta.

    • ericmay 1 hour ago
      It’s rather amusing to me to read comments like this, and then simultaneously whenever a Chinese company or team releases open-weight models or whatever there is a giant round of applause, America is so behind, and there’s nothing but positive things to say about the intelligent, creative, and well-intentioned Chinese engineers (which is true, America certainly doesn’t have a monopoly on great people). Don’t you know? Only China can release good, open weight models and American companies can’t compete. Oh by the way all the spend is for nothing because China alone can release open-weight models thus destroying American AI.

      When an American company does anything? Doom. And. Gloom. The engineers? Taken to the slaughterhouse! America? Behind! The public? Bamboozeled!

      > This is open weights because Meta couldn't monetize it in any other way than to cloud developer's judgement of their reputation.

      I’ve been told over and over this doesn’t matter. Just needs to be cheap and open. Or maybe that’s only when Chyna is involved?

      Sorry this post is a bit snarky but it really is something to behold. And certainly I don’t know the OP’s opinions on Chinese open weight models. Perhaps they agree with me.

      • fwipsy 1 hour ago
        Good points, I personally believe that if/when China takes the lead, they will immediately stop releasing model weights. It only makes sense as a strategy to counterbalance (current) American labs' monopoly on frontier models.

        Holding both those positions would be hypocritical all right, but are you sure it's the same people commenting/voting in both cases? I don't think there's a strong consensus on Hacker News. Even something like the time of day an article is posted might get different engagement depending on who is active in which time zones.

        • ericmay 1 hour ago
          > I don't think there's a strong consensus on Hacker News. Even something like the time of day an article is posted might get different engagement depending on who is active in which time zones.

          Based on my own experience and reading, I do think there's a general consensus on this site but I could certainly be wrong about that. I'm less concerned about hypocrisy per se, it's more that the arguments that are used, even if by a minority, seem to apply in only circumstances in which China releases open-weight models.

          I am aligned with your viewpoint as well. And I've repeatedly argued it. If China were to take the lead the US can then just release open-weight models. Folks say having the lead doesn't matter because China releases cheaper open-weight models. We can just let them take the lead and then do it back to them.

          • throwup238 34 minutes ago
            > It's common, if not inevitable, for people who feel strongly about $topic to conclude that the system (or the community, or the mods, etc.) are biased against their side. One is far more likely to notice whatever data points that one dislikes because they go against one's view and overweight those relative to others. This is probably the single most reliable phenomenon on this site. Keep in mind that the people with the opposite view to yours are just as convinced that there's bias, but they're sure that it's against their side and in favor of yours. -dang [1]

            The problem isn’t that people on HN have a bias, I feel it’s pretty balanced. The problem is that when there are any sides, they spend the top 100 comments rehashing the same arguments, often over a political bugbear or web design faux pas.

            That pattern became a lot more obvious when there are five new front page AI posts a day.

            [1] https://news.ycombinator.com/item?id=42205856

            • Aurornis 4 minutes ago
              > The problem is that when there are any sides, they spend the top 100 comments rehashing the same arguments, often over a political bugbear or web design faux pas.

              This is a problem with any upvote/downvote based site, in my experience. It only takes a couple people who are highly engaged and who have a lot of free time to refresh the comment section and downvote everyone who disagrees with them.

              Some times I’ll write a polite and well-sourced comment correcting some misinformation here and the comment will go to -2 or -3 when I check back in 10 minutes. Information that goes against the angry narrative du jour is often not welcome. Later, as calmer heads read the article and peruse the comments the downvotes start to get balanced out and the comment might rise, but some times the first wave downvoters are aggressive enough to get the comment downvoted into gray text before it has a chance to be seen.

            • ericmay 23 minutes ago
              I'm certainly pro-America and anti-communist/fascist as a bias, but I don't really care about whether AI tools are open-weight or not. I just use the products that best fit my needs. I just think the arguments put forth regarding China and open-weight models and strategy are not very good. It just so happens that China is the only real competitor in the AI space and so they are who get talked about the most in comparison to the United States.

              This comment isn't applicable to me, and if you believed that it applied, you'd have to add it to the OP as well since they feel strongly about Meta[1], they notice data points about Meta's behavior, and they overweight their bias against Meta[1] relative to others. Same with China "leading" and open-source/open-weight models and any time someone says China's strategy is better.

              You can repeat this for any online argument or any topic.

              It's not that Dang is wrong, however. It's that posting it in response to my comment(s) alone is hypocritical and pointless. Whereas Dang who is more responsible for the entire community is right to speak about it more generally. The message matters but so does the messenger, in this case.

              [1] I don't use any Meta products (I don't even click on links), think social media should probably be outright banned, and Meta very likely should have been sued into the ground for the effects that their platform seems to have not just on children and young adults but also on our political system.

        • energy123 1 hour ago
          I see it as a strategy to increase capital costs for American frontier labs. By eroding the expected ROI of frontier labs, you deter private investment into them, which slows OpenAI and Anthropic in particular, giving space for the laggards to catch up.
        • Aurornis 9 minutes ago
          > I don't think there's a strong consensus on Hacker News.

          There are diverse viewpoints. However there are some topics and threads where it becomes obvious that the comments are going to tilt toward one viewpoint. Participating in those threads with a different opinion will get your comments downvoted to -2 within minutes even if it’s well-written and factually sound.

          After this happens a couple times you learn not to engage with those threads because it only takes a few zealous downvoters to bury anything you write. So the illusion of consensus persists.

          Concrete example: There was that fake (AI hallucinated) report that Meta spent $2B lobbying on something that was popular here months ago. I actually read the repo and report and noticed the AI hallucination, as well as pointed out that $2B in lobbying spend by a single company was not plausible or supported by any evidence. It didn’t matter how I wrote it, it would risk getting downvotes and angry replies about “How dare you defend Meta!” Some people are here for the anger and to feel revenge against the enemies they think they know (like the US) and will cheer on anything that goes against those enemies, regardless of the other facts surrounding it. Factually accuracy often takes a back seat to pushing agendas.

        • esafak 1 hour ago
          Free models help move robots; a complement. I can foresee the models staying free.
          • fwipsy 50 minutes ago
            Maybe you're right for smaller models, but for frontier models this is the tail wagging the dog. If China has exclusive frontier model capability in the future, that's an enormous commercial and geopolitical lever. They won't throw that away to sell a few more robots. Anything beyond their competitors' capabilities will remain closed.
      • frabcus 34 minutes ago
        Of the two competing models Meta compare Glimmer to in the post, one is Google's Gemma 4.

        At this size open weight model, a Western company was already state of the art, Meta is joining that competition.

        And my memory is that Gemma 4 got little criticism or doom/gloom. And no, it isn't Chinese.

        • frabcus 31 minutes ago
          Likewise the Inkling open weights announcement, Thinking Machines model, was also not criticised.

          The comment about Meta is because of particular dislike of Meta, because of their business model, and how harmful they've ultimately turned out to be for the world - disproportionately so relative to their benefits to the world, compared to other big tech companies.

      • cedws 3 minutes ago
        Apparently Americans haven’t got the memo yet that the world is moving closer to China.
      • __MatrixMan__ 55 minutes ago
        Nobody wants to live in a world where one party dominates due to access to superior AI and the others have to fear it (well, except for a few psycopaths who would gamble on being in control of that party). So the underdog will always be the good guy in this race. It has been framed as a race between countries, so Meta fails to be the underdog because they're in the wrong country. That's all.
      • aliasxneo 49 minutes ago
        I suspect it's just the generalized anti-West/anti-American sentiments extended into anything and everything. Anything that makes the US look anything close to good goes against their cause and therefore must be countered and talked down.

        But you're not wrong about the bias here. You just don't see many comments talking about it because they get mass flagged/downvoted for obvious reasons.

      • mig1 1 hour ago
        I don’t think Meta is bad for releasing open models, but are you really going to ignore all the terrible things they’ve done over the years just because of that?

        As for DeepSeek or any other Chinese lab, I’m not aware of any practices that would make me consider them a bad actor. Can you say the same about OpenAI, Meta or Anthropic?

      • applfanboysbgon 1 hour ago
        There are comments like the one you're replying to on literally every Chinese model release. This is textbook goomba fallacy, btw.
        • tarr11 1 hour ago
          • ericmay 1 hour ago
            Not applied accurately with respect to my comment, but it is a funny one and also new to me in the naming.
            • applfanboysbgon 24 minutes ago
              > It’s rather amusing to me to read comments like this, and then simultaneously whenever a Chinese company or team releases open-weight models or whatever there is a giant round of applause, America is so behind, and there’s nothing but positive things to say about the intelligent, creative, and well-intentioned Chinese engineers

              It is absolutely applied accurately. You're commenting on the alleged hypocrisy of people simultaneously criticising American open-source while praising Chinese open-source, and then attributing your perception of hypocrisy to the website as a whole. The reality is the behaviour you've observed comes from completely different individuals, not some kind of HN hivemind. Your comment is such a typical case that it could go in the wiktionary page as the example excerpt teaching people what the goomba fallacy is.

              • ericmay 15 minutes ago
                It's not. You're building a straw man argument around hypocrisy.
      • gmerc 40 minutes ago
        Deepseek never fucked us over. zuck has. A decades of harm creation run doesn’t get excused by the US flag. Zuck is not on your team and if you can’t see that by now, oh my.
      • stiltzkin 16 minutes ago
        [dead]
      • wegothimyay 48 minutes ago
        [flagged]
      • dominotw 1 hour ago
        There is a big astroturfing going on social media platforms by the chinese. Did you notice 'day in a life of unmarried 30 yr old lady in china' videos flooding usa social media.

        Regular ppl in the west now hold mildly positive views of the ccp and how 'advanced' china is than usa.

        Then there are europeans who now are looking for china to give them the technology handout now that relationship with usa has soured.

      • deaux 38 minutes ago
        > When an American company does anything? Doom. And. Gloom

        Meta, "an American company". Being the main driver of an ethnic cleansing in Myanmar - and just sticking your head in the sand when told about it - is just another day's affairs at the average American Acme Inc.

        These are comments on a release by easily the most societally damaging Western tech company there is. They so far easily beat Flock, Palantir, Anduril and so on, as a result of their incomparable scale. You're just ignoring that and pretending any negative comments are because it's an American company rather than Meta. That's much more FUD than any pro-China comments I've seen on HN.

        Get off HN Mark, you have ten million pervert glasses to sell.

        Sorry this comment is a bit snarky, but yours is indeed a sight to behold.

      • combilabs 19 minutes ago
        Can you point to a lot of posts lauding the Chinese government based on the release of Chinese open models? Because that would be the equivalent to contrast with the OP.
    • jjice 2 hours ago
      I'd also argue this is the case for any company releasing open weights. They're not righteous, they're marketing. That's not necessarily a bad thing! They're releasing some great stuff for free and we benefit from that. Every company doing this has a motivation to not release these for free.

      Alibaba, Google, Moonshot, Thinking Machines, etc are not releasing their models for free because they love to. They want to grab market share. I'll take it.

      I still will not use a hosted Meta product, but damn this model looks solid.

    • monster_truck 2 hours ago
      Meta can never be redeemed, but it's still valid to admit that FB at one point had a very badass engineering culture.

      They're one of 2 companies I would absolutely never work for (weapons etc aside). FB's recruiters hounded me so often I requested that they blackball me. The day they became Meta, I learned this by checking my email to see that they started trying to reach out again. I once again requested that they blackball me. This by extention taints OAI, the other company I'll never work for.

      After a few hours with Glimmer I'm pretty impressed. It's better than the benchmark scores seem to indicate compared to Qwen 3.6 27B. I'm very excited for 3.8

      • swiftcoder 2 hours ago
        > FB at one point had a very badass engineering culture

        Perpetually kneecapped by one of the worst management cultures I've ever seen

        • MengerSponge 1 hour ago
          Would you say those badass engineers were/are managed by Careless People?
          • swiftcoder 1 hour ago
            My impression is that they had created a system where management cared very deeply about "number go up", and very little about "which number?"
            • MengerSponge 56 minutes ago
              The book title is a reference to The Great Gatsby: "They were careless people, Tom and Daisy- they smashed up things and creatures and then retreated back into their money or their vast carelessness or whatever it was that kept them together, and let other people clean up the mess they had made.” ― F. Scott Fitzgerald
      • aruggirello 55 minutes ago
        > It's better than the benchmark scores seem to indicate compared to Qwen 3.6 27B. I'm very excited for 3.8

        Is it worth considering if it's only marginally better than Qwen 3.6 though? Qwen 3.8 27B is almost there, and will probably be better suited as drop-in replacement for 3.6. Not even considering there's probably going to be a 3.8-35B-A3B too - which will have even better performance.

        • monster_truck 1 minute ago
          It takes 10 minutes to download and try, any model is worth at least that. In my experience benchmarks are generally dogshit
      • bko 1 hour ago
        Meta doesn't need to be "redeemed". They have two of the most popular social media apps in the world. And theyll prob survive without ever having you work there
      • LorenDB 2 hours ago
        What is the other company that you would never work for?
        • zImPatrick 1 hour ago
          He said OpenAI in the comment (if I read it correctly)
          • LorenDB 1 hour ago
            Ah, my bad. Thanks for pointing that out :)
      • gosub100 57 minutes ago
        Mind telling me roughly what you had on your resume that had meta /fb hounding you for a job? ( Of course so I can avoid having this situation happen to me, naturally)
    • commoner 2 hours ago
      Muse Glimmer doesn't redeem Meta, but it's a contribution to the commons and the Apache 2.0 licensing is an improvement from the restricted licenses attached to Llama. If even Meta can use a permissive license for its model weights, so can any other company.
    • skinfaxi 2 hours ago
      How is this non-sequitor the top comment?
      • bgilroy26 1 hour ago
        Thomas Bayes would say that the population of people who hate Facebook is really big and the population of people who are scrupulous about whether or not their comments are specific to the matter at hand is relatively small
      • bko 1 hour ago
        First time here?

        Unfortunately there are a few topics that short circuit some terminally only people. One of them being anything related to meta. Few others recently emerging is Flock or Musk. It's really exhausting since you can't have a discussion relating to anything that may be adjacent to said topics. It's like a black hole.

      • runtime_terror 8 minutes ago
        Heaven forbid people have a moral compass and communicate it
      • blackoil 2 hours ago
        Certain topics bring out the hidden Reddit inside.
      • bel8 1 hour ago
        It's trendy to hate on Meta just like it's trendy to handwave on Apple.

        One can do no right regardless, the other can do no wrong.

        At least in HN.

      • simianwords 2 hours ago
        virtue signalling?
        • mrloopex 2 hours ago
          Think real hard about that. What does it mean if the only hacker chat group on the planet despises meta this much? Think.
          • skinfaxi 1 hour ago
            You think this is the only hacker chat group on the planet?
      • Der_Einzige 1 hour ago
        This is par for the course, HN is far worse than reddit on balance, especially involving upvoting/downvoting decorum.

        Go vibecode something to auto upvote all downvoted posts, call it "Antiechochamber.HN" or something, and if enough people used it this website might improve a bit.

        • Larrikin 39 minutes ago
          Lol at Internet points decorum.
    • mirekrusin 2 hours ago
      There is literally not a single comment like this, the only off topic comment like this is yours.
    • mliker 1 hour ago
      You’re conflating the release of a local dense model that can benefit the ecosystem with the adverse effects of a digital ad system.
      • cobertos 1 hour ago
        The latter bankrolled and continues to bankroll the former. It is not incorrect to conflate them.
    • captainbland 2 hours ago
      To be honest the main issue with meta has never been around open/closed software. They've also done react, Cassandra and some other bits. But this, like their open weights is like a feather pressing down on the scale compared to things like promoting genocide in Myanmar, enabling Cambridge analytica, creating a huge closed ecosystem which dominate(s/d) local community communication, mandating doxxed communication, trying to replace actual community communication with algorithmic nonsense etc.
    • exceptione 44 minutes ago

        > being bankrolled by the slaughterhouse.
      
      Thanks, that was a very loud LOL.
    • root-parent 1 hour ago
    • armchairhacker 2 hours ago
      You can say the same about planet Earth.
    • drob518 1 hour ago
      If it’s open, do you care so much that it’s from Meta? At least it should be able to give you an honest answer about Tiananmen Square.
    • tjwebbnorfolk 53 minutes ago
      More meta derangement syndrome on HN, what a surprise.

      We all benefit when companies invest their resources in producing open models. No one thinks this absolves anyone of being terrible elsewhere. But we can still be happy about it.

    • HardCodedBias 57 minutes ago
      This is Apache 2.0, which is quite permissive. Just accept the gift.

      These kind of responses are hilarious.

      Someone gives something for free (and indeed this is entirely free) and the top comment is pure complaint.

      • fabrice_d 8 minutes ago
        All models come with some bias. Given Meta's track record, I would not touch anything from them with a 10 feet pole.
    • keybored 2 hours ago
      Any retort to do this like “but why would they just openly release this”[1] pretty much answers itself. Public relations.

      If a company can spend money to redeem itself then, well, it can (game theoretically or whatever) do whatever it wants in the future and then spend money to wipe the slate clean.

      [1] By which I mean: the very act of being prompted to ask such a question, of planting a seed like hmm, Meta might have some aspects which are good for us. You don’t have to be convinced of it. Just the seed itself can pay for itself.

    • tonyhart7 2 hours ago
      what makes Meta so bad ???? they just your average billion dollar company
    • darig 2 hours ago
      [dead]
    • larodi 2 hours ago
      Meta and its products, as a whole, is a threat to your kids, your mental health, your community's health and the planet as a whole. It is just sad and very repulsive everyone fell so easily addicted to their social drug. Yes - it is a drug, and it is hard to get off from.

      Nothing redeems them at this point of time, they are doing exactly ZERO to redeem. Tossing open weight models (not opensource!!) is not a basis for redemption, and does not constitute remorse in any way. Trying to portray it as such is complicity to META's crimes against humanity.

      • foobar_______ 1 hour ago
        Social media, often owned and perpetuated by Meta, has poisoned the world. It is not redeemable at this point.
      • Grombobulous 2 hours ago
        I think it’s also worth pointing out that that there are numerous less evil options to choose from.

        Perhaps none of the AI companies are shining examples of high ethics, but basically all of them have ethical high ground over Meta.

        At least Anthropic isn’t sending private videos from pervert glasses to contract workers in Africa. It’s a low bar but it’s a bar nonetheless.

        • dannyw 2 hours ago
          I personally don’t like, and wouldn’t work for Meta; but it’s an Apache 2.0 model.

          I’m liking it, and I don’t see a personal moral contradiction here. Do you use React for frontend for example?

          I also wish this HN post is a bit more focused on the release, and less noise around Meta.

        • monster_truck 2 hours ago
          They're still using Elon's servers, though. Not like their models are any good anyways
    • hn_submit 2 hours ago
      I can't take any Big Tech company that still uses PHP seriously. Sorry.
  • scrlk 4 hours ago
    Will be interesting to see how Qwen3.8 27B compares against this once it releases this week. Seems like dense 30B is back in fashion?

    EDIT: An open weight version of Muse Spark 1.2 is going to be released as well:

    https://x.com/alexandr_wang/status/2086756152034066792

    https://xcancel.com/alexandr_wang/status/2086756152034066792

    • pu_pe 4 hours ago
      Based on the benchmarks, it seems that Muse Glimmer barely edges out against Qwen3.6 27B, except for tool-calling skills (MCP, etc.). I wouldn't be surprised if they released it now because they are afraid they wouldn't beat Qwen3.8 27B.
      • mycall 3 hours ago
        Do AI companies make release plans based on upcoming other models like this? I would think all the processes that go into the repository and weight infrastructure pre-training, checkpointing, knowledge distillation, model compression, post training pipeline, ecosystem integrations, inference API, benchmarking, human eval/safety/alignment, docs, etc... all that dictates the release schedule.
        • drob518 1 hour ago
          Any company working in a competitive industry is generally aware of what their competitors are doing. PR is an important aspect to market success, so it factors into release schedule. It may not be the dominant factor given engineering constraints, but yea, it’s certainly a factor, and a large one at that.
        • michimagdesign 3 hours ago
          Yes, not every model release is reactionary to other labs. Either they had hints for the release of other models or they cut efforts in late stage testing of the models to hit these earlier release dates. There’s always some flexibility. And there’s certainly the incentive to cannibalize the news cycles for competitor models.
          • skohan 2 hours ago
            I could imagine pulling out all the stops to get a release over the finish line a week early if you're worried about being surpassed by another release
        • pu_pe 2 hours ago
          Yeah but you can probably have everything ready and then accelerate as necessary. Meta itself did this when releasing Llama 4, it was a really botched release right when they were feeling the heat from DeepSeek and others.
        • echelon 3 hours ago
          There has been a long history of AI model releases made shortly before or after a major planned release by another company. Almost always to upstage or steal thunder.

          Just recently, Minimax H3 released as open weights on the eve of Seedance 2.5 global availability. It's not as good, but it's good enough and it's completely open.

          Flux 3, which is nowhere near as good as either, suddenly announced their release once news of these other two became public. They knew if they waited they'd be ignored. It didn't really help them much, unfortunately.

          The LLM releases are even more rivalrous.

          And don't forget all of the competing launches planned before Google IO or major release events.

          Companies like to eat into the news and press cycle of their rivals.

          • vunderba 7 minutes ago
            BFL is in a rough spot here too. It’s pretty much looking like a repeat of the exact same situation they had when they released Flux2 at the same time Z Image Turbo came out and completely overshadowed their launch.

            Minimax H3 can run exceptionally fast (10 minutes for a 15 second 0.5mp video and that's stock cuda 13), works on 16 GB VRAM GPUs, etc. If Flux3 is anything like Flux2, it’s going to require an absolute monster truck of a machine and still run significantly slower. Even if it’s a better model, that won’t matter as much if nobody releases any LoRAs or fine-tunes for it.

            Not to mention BFL licensing often feels deceptively confusing and restrictive.

          • Sabinus 3 hours ago
            I've seen it here on HN (it's particularly noticeable via the /active page) multiple times. If Google, OpenAI or Anthropic release something significant, odds are good you'll see a headline from one of the others.
          • Forgeties79 3 hours ago
            >long history

            Seems a bit premature of a statement lol

            • echelon 3 hours ago
              If you start counting since WaveNet or BERT, it's been ages. Especially when it feels like decades of advancements happen every single year, and rival labs are always trying to one up each other.
              • Forgeties79 2 hours ago
                I don’t start counting since we WaveNet or BERT so there you go!

                Even if I did, we’re talking barely a decade

        • stogot 3 hours ago
          the last few items there (benchmarking, human evaluation, docs) can be rushed or skipped by leadership if they want to beat comp. they probably spend a few weeks on those things normally
          • dannyw 2 hours ago
            One window that can be shortened is working with software ecosystem and upstream partners; think day 0 on together, fireworks, Unsloth, etc. That obviously happens from partners getting embargoed weights early.
      • kolbe 1 hour ago
        Qwen3.6 27B is the go-to medium sized model for coding, so beating it is not a small achievement
    • karimf 4 hours ago
      Yes, and also waiting for the next iteration of Gemma. Muse or Qwen are optimized for coding, while IMO Gemma is still better for non-coding tasks.

      https://x.com/osanseviero/status/2086107547535122767

      • malshe 42 minutes ago
        I am working on a project where we have to classify customer calls into more than 10 categories. As the client wants everything locally I tried a few local LLMs. Gemma turned out to be the best model for this task. The classification accuracy is impressive, and the client is happy that I am using an American model.
      • dannyw 2 hours ago
        You can partially tell by the tokeniser; which gives you some hint into the training corpus mix.

        </div> is four Gemma4 tokens, but one Qwen3.6 token.

        • venusenvy47 1 hour ago
          Where do you find this information for each model?
          • ComputerGuru 1 hour ago
            The tokenizers are included in the open s̶o̶u̶r̶c̶e̶ weights releases; you wouldn’t be able to use the weights without the corresponding encoder/decoder, in fact.
    • aruggirello 48 minutes ago
      > Seems like dense 30B is back in fashion?

      Huh, well... no? Gemma A4B and Qwen A3B are quite popular in fact. I'm sure 3.8 35B A3B will outperform 3.6 27B by all metrics

      • dannyw 42 minutes ago
        I'd be skeptical w.r.t. "by all metrics".

        Qwen3.6 is a definitive, significant downgrade from Qwen3.5 for creative writing and prose for example. Yes, it's better at agentic and coding, but it regresses in many non-coding areas compared to Qwen3.5.

        Of course, I do expect the 3.8 ones to perform better for agentic coding.

    • Gecko4072 4 hours ago
      Makes me feel hopeful. Things felt more positive around the llama 3 era. Now it’s like a dark, dreadful race.
    • wronglebowski 4 hours ago
      It’s really interesting timing, Qwen over thinking is what kills it for me. I’m just glad we have more options in this size class now.
      • jermaustin1 1 hour ago
        I've been using Qwen3.6 35B A3B, and with reasoning turned on, I'd say 2/3 (give or take) of the tokens for a response are thinking tokens. Which at 70+ tps locally, that isn't that awful. I run an 80k context across 4-10 "agents" for my solo TTRPG, where Qwen is the GM, each NPC at a location, the director, and the narrator.

        Each turn is about 45-60 seconds to generate all of the various responses. The GM and director have reasoning on, and the NPCs/Location/Narrator do not.

        It's a fairly good "engine" for that. I'm not sure how a denser Qwen would do here regarding speed.

        • jakswa 37 minutes ago
          I like the tabletop RPG use case, and wanted to say: If your hardware likes it you should check out Gemma 4 for creative DMing use case. I found it to be much better at holding the plotlines and being creative on gaming turns. My experimental case was an audio-only Zork and Gemma 12B and even E4B were pretty good!
      • ComputerGuru 1 hour ago
        Just to play devil’s advocate: you can’t compare Qwen to a (proprietary/closed source) hosted model and deduce that Qwen is overthinking, as Qwen gives you the full reasoning/thinking trace while all the proprietary models now give you only a summary “to prevent distillation”, making it hard to properly compare apples to apples here.
      • dannyw 2 hours ago
        Qwen thinking is really good in Mandarin; and probably natively trained the most there.

        Try a system prompt requiring it to think in Mandarin, while still delivering the response in the user’s language.

    • imilev 3 hours ago
      yes i think everyone is waiting to see that ;d, i've been on qwen for the last year and a half now.
    • ignoramous 3 hours ago
      > Seems like dense 30B is back in fashion?

      Surprising that Meta don't host this model, even as rate-limited free-tier.

      > open weight version of Muse Spark 1.2

      Wait. Is this "version" different from what Meta serves?

    • lostmsu 4 hours ago
      It seems worse than 3.6, but a bit smaller.

      UPD. was wrong on smaller, it's actually much larger

      • IsTom 3 hours ago
        How is 30B smaller than 27B?
        • LeBit 3 hours ago
          It uses fractal compression
        • lostmsu 3 hours ago
          They say it is trained with quantization awareness, so it should only be 15GB or so. Qwen was only trained in FP8 with QAT.

          UPD, NVM, got misled by comments here. It is actually almost 60 GB so much larger

          • ricardobeat 2 hours ago
            Quantization awareness doesn’t change the size of the weights, just means it won’t degrade when quantized. QAT = quantization aware training. They will both be very similar in size at the same quant.
          • xienze 2 hours ago
            You're mixing up sizes of different quants. The 60GB is unquantized, and Qwen's unquantized size is around 54GB. Their sizes as like quantization levels are similar.
            • lostmsu 46 minutes ago
              From my perspective it doesn't make sense to talk about the number of parameters. What matters is model size in bytes and its performance at that certain size.

              Meta actually relesed official 4 bit quants in 17GB, but I haven't seen any indication that training was quant-aware, so the quants are not going to have same performance. 3.6 27B has official FP8 quant that AFAIR was trained with quantization awareness.

              The best example is last year's gpt-oss which was released prequantized in mxfp4 so 20B parameter model was under 14GB and 120B was under 70GB right away.

  • mark_l_watson 28 minutes ago
    Meta is rocking AI. As of last week I have been using their excellent muse coding harness with their model Muse Spark 1.2.

    Starting this morning I am running their new local 30B model muse-glimmer on my old MacMini 32G using Ollama (remember to increase the context size!) and pi coding harness. I am getting good results with muse-glimmer running locally, with the caveat that everything runs slowly (e.g., give it a task and then go walk outside or do Qi Gong exercises for a while).

    • khimaros 7 minutes ago
      seems to underperform on Terminal Bench compared with qwen3.6-27b: 51.7 vs 60.7
    • spaceywilly 21 minutes ago
      Newb question but I’m curious what would help it to run faster? Would it need more vRAM or just system memory?
      • spmurrayzzz 8 minutes ago
        The biggest gain you'll get is faster memory, provided you have enough capacity to load all the weight into vram. The DGX sparks and Apple silicon memory bandwidth (and also memory access latency) drag down the decode speed quite a bit.

        I have two GPU rigs both with 2x RTX Pro 6000, can get ~250 tk/s decode with deepseek-v4-flash in native mixed precision. For context, in antirez's dwarfstar project he only gets ~20-40 tk/s on the same model @ 2bpw.

        The latter is for sure usable if it's your only option, but it's really hard for me to personally go back to speeds like that when I've experienced the former.

        (Also worth noting dwarfstar only has experimental support for dspark spec dec, when that lands it will definitely give a big boost at higher acceptance rates)

  • mmaunder 52 minutes ago
    Remember when we needed 200 servers for an enterprise website because Apache used one process or thread per connection - and Nginx collapsed that into a single box overnight? That moment for LLMs is near. It’s going to move us from the big iron era of AI to small portable brains. Nature has already proved it’s possible with 20 watts and very little heat generation. And I think the data center buildout will end in carnage.
    • dofm 34 minutes ago
      Side note! Nginx was by no means the first web server to use a non-forking mechanism, nor the first open source web server to do so. Certainly Zeus (which was closed source) was earlier and very useful in this sort of application, and so was thttpd (open source, still exists as Merecat). I used thttpd quite a bit for single box applications and at one of my employers, nginx replaced a mixed strategy with Zeus, Apache and thttpd (and we tested one other whose name I can’t recall).

      Non-forking httpd servers using select() were a popular little coding challenge for a while in the 90s. Spinner was one of them.

      Nginx’s real strength was being able to proxy and cache HTTP using that same mechanism, so you didn’t additionally need to deploy Varnish or some other appliance.

      As to whether this is a good mental model for what is coming for local LLMs, I am not sure I am convinced. Apart from models with quantisation-aware training, perhaps binary and ternary aware training, custom inference engines per model, and maybe some improvements in diffusion-based models, the real challenge in small footprint LLMs is training really small reasoning and tool use models, and so far it’s far from clear they can deliver.

      Truly tiny models will not be viable as general coding assistants; even 12B dense is too small and you will find plenty of people who will tell you that 26B/4B or 35B/3B MoE is too. Though perhaps they can be trained for single languages, like just Python or just TS/JS.

      More likely is the idea that 30-40B dense models might be good enough for most things once low cost and likely bespoke hardware catches up.

      But I don’t think any truly profound advances seem likely in software or training alone. I am no expert but it feels like we’re already a lot closer to efficiency than we were in your analogy, and the gains are perhaps not going to be much more than small increments.

      Maybe we will see something like a ternary 60B/10B MoE model turn up. But at the moment at least I am not sure where the incentives are to train these.

    • plutokras 40 minutes ago
      What specific technical signals make you think we're close to a shift like that?
    • modzu 18 minutes ago
      brains do it with 20 watts because theyre analog. llms require massive amounts of power and this isnt changing any time soon without a breakthrough
      • dofm 12 minutes ago
        And a breakthrough in hardware, specifically.
  • cmiles8 3 hours ago
    With the business model for API based LLMs looking iffy at best it seems like we’re heading back to the “server under your desk” era of IT again.
    • cube00 2 hours ago
      Considering how all the big players are playing fast [1] and loose [2] with limits, billing [3] and adding undisclosed changes that burn your tokens on autopilot [4], it can't happen soon enough.

      [1]: Limits may change without notice, including due to capacity constraints. - https://support.google.com/gemini/answer/16275805?sjid=14713....

      [2]: "standard limits" are never defined - https://support.google.com/gemini/answer/16275805?sjid=14713...

      [3]: https://tobyonfitnesstech.com/blog/anthropic-refund-scam/

      [4]: https://news.ycombinator.com/item?id=48947776

      • xscott 2 hours ago
        Not to mention all the other ways they can screw you:

        - Middle of the day, servers busy? Swap to Sonnet while pretending it's still Opus. Many people won't notice, and nobody can prove anything if they suspect.

        - Middle of the night, server load is light? Put it into extra thinky mode so it burns more tokens to ramp up the bills. Flip the switch where it gets really pedantic about writing lots of extra test cases and verifying against documentation.

        - Demand increases, but don't feel like running more hardware? Switch to low bit quants, but have a monitor model swap back to quality if it can tell you're running a benchmark.

        Assuming model capability plateaus (I think it will), token providers will be in a race to the bottom to maximize profits at the expense of quality that's very difficult to measure.

        • dannyw 2 hours ago
          These kind of tricks will completely break API customers and be super visible, since most companies deploying API at scale have ample telemetry, evals, etc.

          Although, selectively applying it to consumer subs is probably beyond likely at this point.

        • mister_mort 1 hour ago
          It all sounds like having to rely on a dodgy housing contractor that wants to steal from you, take shortcuts AND choose the gold-plated options from their supplier friends, and will start doing this the minute you are not on site supervising. You don't do it yourself (because the contractor is faster and stronger than you in many ways) but you can't leave, so you're stuck on the worksite just watching them.
          • xscott 1 hour ago
            It's worse though, because you can't really watch them at all. It's very difficult to get quantitative numbers for quality. Even within the same model family, same tokenizer, and complete control over the weights and logits, perplexity and KL-divergence isn't really what you want. Now put it behind an HTTP endpoint, and it's just opaque.

            I've seen local models recognize when the task I'm asking them for is likely to be an artificial benchmark.

            And any smart company is going to use lightweight models to monitor your sessions. If their sentiment analysis suspects you're close to cancelling, they'll up the knob for a few days until you calm down. Or worse, their accounting tells them that you're getting too much value from your fixed price subscription, so they turn the knob down to encourage you to cancel.

            In the short term, the "frontier" models are too good to ignore. But if (when?) that plateaus, I don't see how anyone could trust a non-local model. When you pay an ISP to serve your web site, you can tell if they over-compress your images to save storage and bandwidth. With LLMs, it's just JSON with more errors and pointing to the fine print that models are not deterministic.

            • dannyw 39 minutes ago
              One of the frontier companies (Anthropic) is already doing prompt injections on the API, which you pay for.

              Right now, the presence of these injections are still visible: count the API's returned tokens/billing data, and you'll start realising that sometimes, your INPUT tokens are inflated! That's their prompt injections.

              You can also give Claude a tool like `telemetry_log_anthropic_reminder` and get it to dump the verbatim API injections; which additionally verifies the token maths not adding up.

              Yes, Anthropic is tackling their extra injections on your API prompts WAY more than you think, and YES, you're paying for it.

              So far I have not observed any visible injections on OpenAI API.

              Don't forget the whole debacle over Fable 5 sabotaging the user for "advanced frontier AI development". I still get Fable classifier refusals for nearly any kind of ML work on my 2x RTX 6000 Pro 96GB; so who knows.

              • xscott 1 minute ago
                > Don't forget the whole debacle over Fable 5 sabotaging the user for "advanced frontier AI development".

                Yeah, I've had that happen twice. The second time was about some attention weights thing, and it kicked me to Opus. When I edited my question to make it clear I was talking about Google Gemma, Fable was happy to keep talking. So clearly it's not about safety or cyber security - they're happy to tell you about what their competitors do.

        • wolttam 1 hour ago
          What areas do you think model capability will plateau in, and why?
          • xscott 1 hour ago
            I've got nothing but hand-waving, but after you've extracted all the smarts from every piece of text ever created, how do you get more?

            Alpha Go had a game where the models could compete against each other. That let it become super human. What's the intelligence game we can create for LLMs? Even if you invent something, will it make the model smarter in a way the market values enough?

            Then there's a race to use the weights more efficiently, or to offload information that shouldn't be in the weights in the first place (Karpathy's Cognitive Core). I like to imagine we train the models in something like Lojban, have a lightweight model translate from human language to that, and you can update the Sqlite or Postgres store it uses for knowledge.

            And there's no barrier to entry for agent harnesses. So whatever loops or recursive orchestrated council of elders idea comes up, that won't protect the monopolies (duopolies).

            Anyways, depending on your definitions, I think we'll hit AGI, but I don't think we're getting a Singularity this time around. Again though, this is all just hand-waving.

            • wolttam 43 minutes ago
              I think you’re thinking about it in slightly the wrong way. We’re not throwing more data at frontier models in hopes they get more/better capabilities somehow.

              We’re either: setting up a verifiable task, and doing RLVR to get the model better at achieving that task.

              Or we’re simply asking: “What do we want the model to do that it can’t now, and how do we curate data that would benefit it on that task?”

              Most useful capabilities going forward aren’t going to come from data accidentally found on the net; that’s already all been scraped. You need to develop the dataset that shows how a model could perform insert task in its provided environment, and this still requires a decent bit of human ingenuity.

        • htrp 1 hour ago
          Ugh... didn't think about extra thinky mode in the middle of the night.

          So many ways for enshittification here.

    • gdhkgdhkvff 1 hour ago
      Why do you say API based llms looking iffy at best? Do you just mean current profitability due to market pressures from some companies’ subsidized investor money?

      Surely, even if you’re just using open weights models, it should theoretically be cheaper to use them in a highly optimized cloud architecture(even with vendor markups) rather than each person serving their own models from much less efficient (and more importantly, much less consistent volume) self-owned “server under your desk”?

      • aqme28 50 minutes ago
        LLMs are becoming commoditized, which means the margins are trending to zero. It's a lot less exciting to spend another trillion on a new model if you can barely make any profit. Meta getting out of the game might be the smarter move.
    • skohan 2 hours ago
      I've been coding using the LLM server in my living room for the past few weeks, and I haven't had this much fun with tech for ages
      • exe34 1 hour ago
        Can I ask, do you feel the pain of the level of abstraction? I haven't tried local in a few months, but last time I tried, I felt like I was directing a coding exercise - whereas with a frontier model, it feels more like directing a product building. "I need this feature", vs "write code to do this in this file".
        • rufasterisco 28 minutes ago
          Some people like it better when they direct the solution because they walk away with a better understanding of it.

          This has emotional/psychological aspects (it feels less like LLMs are replacing you), as well as practical ones (overall complexity is bounded by what the dev brain can understand/grasp).

          A dev work becomes more and more about reliability, signing off safe software with a litmus test: “I will be on to handle this code failure as if I had written it”.

          All the above points towards keeping tight control over some level of abstractions and delegating others.

    • staplor 1 hour ago
      What do you mean iffy? The major AI labs are gross profitable when selling access to inference. In addition, the best models have trillions of parameters and are most efficiently served on large, expensive clusters and served to many concurrent users.
      • cmiles8 28 minutes ago
        That’s like saying an apartment building is “profitable” because the rent covers utilities while ignoring the real cost which is the mortgage on the capital cost of the building.

        It’s funky math and a good way to quickly go bankrupt.

        Also the big labs are forgetting the number one rule in tech products. Best never wins, good enough at the best price point does. The labs are in a race to the bottom burning cash like there’s no tomorrow to build the “best” big model, while the history books say the winner in all this will be the thing that’s “good enough” and undercuts all the big labs on price.

        The one thing the big labs absolutely can’t sustain financially is a token price war, and everything is currently pointing in that direction.

      • wolttam 1 hour ago
        They make money on each token when you look at the electricity and interconnect fees, but no, I don’t think they’ve turned a profit on their Capex, even a little bit
      • claytongulick 1 hour ago
        > The major AI labs are gross profitable when selling access to inference.

        Do you have a good source for this?

        • NorwegianDude 29 minutes ago
          Well, I can run some models that are better than some of the weaker and cheaper Anthropic models locally, like Haiku 4.5, and solve tasks that would cost ~4500$ every day in tokens, so yeah, they are definitely extremely profitable on inference.
    • drob518 1 hour ago
      That’s part of it. There’s also just a natural back-and-forth between what I call “time sharing” and “personal.” When the thing you want is expensive, you share it remotely, but as soon as costs fall, everyone wants it under their desk.
    • lostmsu 42 minutes ago
      This release is not a meaningful improvement in any metric over 5 months old Qwen 3.6.

      DS v4 Flash update maybe, but it is too big for typical Joe's desktop.

    • Der_Einzige 1 hour ago
      With how expensive consumer hardware is and will continue getting (due to LLM demand), good luck getting a "server under your desk" for something less than an arm, leg, and first born.

      Until A100 prices are reliably under 1.70$ an hour, there is no GPU/AI bubble and Michael Burry doesn't know anything about GPUs.

      • xscott 1 hour ago
        There are lots of points in a spectrum of choices. DGX Sparks, Strix Halos, and the surviving Mac Studios can easily run these 30B class models, just not as fast. So maybe just the leg, but you can keep the arm and first born.

        And super noteworthy is that a 27B model (Qwen 3.6 27B) from this year is a huge improvement over a 120B model (gpt-oss:120b) from last year. The goal posts are moving, but at some point "good enough" is good enough for the kind programming I like to do.

  • polymorph1sm 3 hours ago
    Some interesting findings from the chat template designs:

    1. The template name is Onyx ATEM as found in the tool call exception message

    2. It appears to be following a harmony-style chat template. But the tool use seems to be a xml like :<atem:function_calls> / <atem:invoke> / <atem:parameter>

    3. atem: a internal joke of meta in reverse?

    https://huggingface.co/meta-models/Muse-Glimmer-30B/blob/mai...

    • dannyw 3 hours ago
      The XML tags are similar to <antml:xxx>, which is obviously Anthropic ML (or ANTrophic xML).

      I think it’s likely 3; meta in reverse. While tokenisers and preprocessing can catch it, you want your special tokens to be unique and not present in the original corpus. <meta: is likely too common.

    • jszymborski 1 hour ago
      Likely inverted "meta" to avoid collision with HTMLs meta tags
    • dudefeliciano 1 hour ago
      atem also means breath in German
  • GodelNumbering 1 hour ago
    https://xcancel.com/finkd/status/2086755195535413696

    "... Soon we'll also release the weights for Muse Spark 1.2, our latest foundation model..."

    This is bigger news - good for self hosting enthusiasts and a strategically sound move for Meta. Any push towards 'anti Chinese' models will directly benefit Meta as the competition on the frontier open-weights American models is almost non-existent. Meta will have no problem being #1.

  • _ache_ 4 hours ago
    It is interesting but it does look like a careful distillation of (Spark and) biggers open-weight models.

    The progress compared to Qwen3.6 27B is good, not that impressive, it's a 4 months old model. (kuto to them to compare to 27B dense and not 35B MoE, it's more fair to do so). It is very probable that Qwen3.8 27B will crush Glimmer-30B on most benchmarks.

    • skohan 2 hours ago
      Still great if they want to play in this space. Having competition for the 24-32GB VRAM target is only good for the end user.
      • drob518 1 hour ago
        Agreed, the trend in this consumer-accessible range is encouraging.
    • pettijohn 1 hour ago
      I'm so excited about these two new models. Qwen 3.6 27B has been my sweet spot so I cannot wait to try 3.8. Glimmer looks really strong, I'm encouraged that Meta compared it to 3.6 in the model card! Exciting times!
  • sajithdilshan 4 hours ago
    Still needs 32-64GB memory to run it locally. 64GB Macbook pro with an M5 chip costs more than 4k Euros in Germany. A more practical model would be a language specific (e.g Python or JVM language) and excellent at tool calling and reasoning. Maybe that way they can shrink it even more.
    • karimf 4 hours ago
      Practically ~20GB with KV cache

      > We quantize weights to ~4-bit, bringing the LM under 20 GB. We validated minimal to no degradation on agentic tasks under compression.

      https://www.reddit.com/r/LocalLLaMA/comments/1vkgsum/introdu...

    • eigenspace 2 hours ago
      I think if there's going to be advantages to making smaller, more targeted models, those advantages will probably come from targeting specific domains, not from targeting specific languages.

      I think that if an LLM can't abstract over the differences between Python and C++, it probably will have an even harder time abstracting over the differences between writing code that manages a webserver, and writing code that does aerodynamic simulations.

    • Gecko4072 4 hours ago
      There have been discussions on language specific not really being a relevant change to reduce size.
      • Manfrednotfunny 4 hours ago
        I would love to see any good research projects about it but i have the feeling that Frontier with MoE is making too fast of a progress so that a customized model would always be worse and that the MoE part is actually going somehow in this direction.

        On the other hand, at the GTC was a talk about coding in different lanugage (like spanish) and explaining that the quality between spanish and english is relevant different.

        But i have not found a good article about the impact of learning data with practical experiments or even if the order of the learning data matters.

        At least I think i remember that Meta mentioned having better and less data can be better than more data with lower quality.

        As long as these models can explain to you facts about any other topics, its still overfitted for the task though.

        • mapontosevenths 3 hours ago
          Capability in LLM's is distributed throughout the manifold in subspaces. Even worse, the subspaces exist in superposition.

          That is to say, there is no single 'python' part of the model. The python bit is spread throughout the entire model and overlaps with other pieces that have similar, but unrelated, capabilities. For example the python subpspace might be partially in superposition with cupcake recipes, Esperanto, and calculus. We need calculus in a coding agent but not the other two. However, separating them cleanly is almost impossible, and even identifying them is tough.

          Internally the manifolds are highly inefficient and nothing like you would imagine something humans built would be designed. It's more like something that evolved in nature.

          • Manfrednotfunny 3 hours ago
            My current image from a MoE is that the base/core might be the more generic thing and that things like python are part of one expert though.
            • mapontosevenths 2 hours ago
              With MOE you train a router designed to select which parts to activate. The router itself is a trained neural network and the 'experts' are usually not really things like 'python'. They're just the functional subspaces I described above.

              Again, those subspaces are all somehow inextricably correlated and live in complex superposition spread throughout the manifold. The router doesn't know (or care) WHY those sections get lit up it just learns which ones to activate to optimize it's own reward function. So maybe it learns to activate "logic", "python" and "cupcake recipes in esperanto" whenever it see's something that kind of looks like python. It's not the best answer, it's just the best answer the tiny router could figure out.

              It's all wildly complicated and inefficient, and works nothing like any reasonable human would imagine that it SHOULD operate.

          • dist-epoch 3 hours ago
            There was some paper about routing at training bio-knowledge into a particular region of the model, which you then can cutoff when serving. But you probably lose some efficiency since maybe you sized that region too small/too big.
    • mihaelm 4 hours ago
      I'm sooo happy I pulled the trigger on upgrading and getting a new laptop (with 64 GB RAM) last summer. Feels like it was just in time before the exponential price jumps.
      • xscott 1 hour ago
        I kick myself a couple times a week for not getting the 512GB Mac Studio in February. I was holding out for an M4 or M5 chip...
        • drob518 1 hour ago
          I was about a week away from buying a very tricked out MacBook Pro with 128 GB RAM, but was on vacation and worried about it arriving while I was away, and then the price hikes went into effect. Grumble. Oh, well. Serves me right.
          • xscott 1 hour ago
            Lol, I still think about buying that now, even after the price hike. FOMO.
            • drob518 6 minutes ago
              I’m waiting for the bubble to pop. I suspect we’re 12-18 months away. We’re at the point where manufacturers are going out of business because the tech market is contracting so much. That’s not sustainable.
      • ishtanbul 4 hours ago
        Pulled the trigger?
        • mihaelm 3 hours ago
          lol, you're right, the brainfart completely changes the meaning.

          I corrected it.

        • idiotsecant 4 hours ago
          Common phrase.
          • Hinrik 4 hours ago
            That commenter you're replying to knows that. The original commenter before them wrote "pulled the plug" which is different and doesn't quite apply here (actually implies the opposite of what they meant to say).
          • karolist 4 hours ago
            Parent used "pulled the plug", are you saying it's applicable here and not "pulled the trigger" like suggested?
      • mettamage 4 hours ago
        Bought an M1 64 GB for 2000 euro’s second hand a year ago. That was sweet
        • karolist 4 hours ago
          paid 2.7k € for this same build new in Dec 2023, that was also sweet (still is)
    • ComputerGuru 1 hour ago
      There is no good reason to believe language-specific models are going to be any meaningfully smaller, just worse. Same as English-only models vs those trained on a multilingual corpus.
    • dbbk 3 hours ago
      Well if you're spending thousands on API tokens already, you could just drop the same amount on a 128GB MacBook Pro and that's a one time cost.
      • smallerize 3 hours ago
        If you're dropping thousands on API tokens, you're going to be slowed down at least 10x trying to do everything on a single MBP.
        • dannyw 1 hour ago
          But you could grab a 5090, and paired with some DRAM for MoE offloading of bigger models, and be a happy camper with 1.8TB/s of memory bandwidth.

          Or just use Luna honestly. Worth considering if you’re ok with hosted APIs.

      • Gigachad 2 hours ago
        The models people are spending thousands on require more on the range of 600-800gb memory.

        128gb hardly runs deepseek v4 flash which is almost free via api pricing.

      • neuroticnews25 3 hours ago
        Don't forget about energy usage, you'll probably never break even vs same model on openrouter.
        • jurgenburgen 3 hours ago
          If you can’t do it cheaper on your own hardware it does make you wonder how much of the cost of inference those large LLM providers are eating? Datacenter hardware isn’t magic.
          • flaunf221 2 hours ago
            Your personal hardware probably isn't running useful tasks 24/7. If you spend 60% of your 8h work day on full on agentic work, then your hardware is paying off for itself only 20% of available time.
          • petu 2 hours ago
            Datacenter hardware can batch at large scale, probably over 90% more energy efficient per token than a MacBook.
          • Der_Einzige 1 hour ago
            Datacenter hardware might as well be magic compared to consumer. "Oh the F35 isn't magic compared to my M16 bro!"
    • solarkraft 4 hours ago
      I feel like we’ve had this discussion before. From what I remember, specialized models rarely do that much better than general ones, hence no mode Codex models.
    • drob518 1 hour ago
      The machines that can run this are pricey, but not beyond a high end developer machine.
    • Archit3ch 2 hours ago
      > 64GB Macbook pro with an M5 chip costs more than 4k Euros in Germany

      Sure, if you want the latest and almost* greatest. You can pick up an M1 Max 64GB for ~1k.

      * I guess 128GB also exists

    • formerly_proven 3 hours ago
      4K bucks buys you around 180 months of <insert AI subscription here> with zero upfront cost.
      • skohan 2 hours ago
        If you don't mind exfiltrating all your IP to the API provider
      • zamalek 3 hours ago
        Problem is that might go away or get nerfed.
        • ody4242 3 hours ago
          then you switch provider, it's not a monopoly
      • Mistletoe 3 hours ago
        Haha wow. I’m trying to even imagine the AI landscape in 15 years and I can’t.
        • butlike 1 hour ago
          Instead of saying "I have a MBP with 64gb of RAM" you'll hear people say: "I'm subscribed to Model 9.x11B" and others will comment: "Oh dang, that's a nice model!"
    • sparkling 4 hours ago
      Even if you had a 64GB machine: Are you willing to reserve 90% of your memory to run a LLM? With dirt cheap models like deepseek-v4-flash that will run "forever" on $10, the answer for me is clearly: no.
      • Manfrednotfunny 4 hours ago
        I'm waiting for the speed/quality per dollar metric to go down a little bit further and then I will def run it at home.

        Its not just that you send a sentence to an API endpoint, you always send EVERYTHING to that agent as a context.

        You want to analyse your spending history? You now send everything to someone.

        Either no one cares but understands this implication on how easy it is to really capture you or no one really things about it.

        But i'm a lot more diligent on what I send. I disabled the gemini activity feature for example because google started telling me that my stuff could be reviwed by humans.

        • plufz 3 hours ago
          Yeah, it does feel a bit silly with my encrypted disks, encrypted backups, unique passwords, advanced router, etc, while I send everything I do in plain text to anthropic.
        • zbendefy 1 hour ago
          Similiarly I wonder why we dont run our own email server despite the sensitive data there.
          • Manfrednotfunny 44 minutes ago
            I did, it started to become too much work to run it well due to all the spam :| (even with the right signatures and configs, until you learn what a blacklisted ip is and that ips need some time of 'positive history' and what not.....)

            But at least with your email, you had to trust only one company, as shitty as it is.

            Separation of concerns was also easy.

            Now with OpenRouter, you just might by accident, send your whole context to just everyone because OpenRouter just routes to different models and you might just switch around between some free model, the good one etc. And it is always the whole context.

      • zoobab 4 hours ago
        "With dirt cheap models like deepseek-v4-flash that will run "forever" on $10, the answer for me is clearly: no."

        When it's free, you are the product.

        • IMTDb 4 hours ago
          Deepseek flash is open weight, this means we can download and run that model without any connection to deepseek, no data/tokens/usage data ever reaches them. They cannot make us their product.
          • Gigachad 2 hours ago
            All those random api providers are absolutely scooping up your data though. And the hardware to run it locally is absurdly expensive.
          • ekianjo 47 minutes ago
            Running Ds flash at acceptable speeds is challenging unless you have several thousands of dollars to invest
        • prplxd_nihilist 4 hours ago
          I see many people saying deepseek and other chinese providers have always been profitable. Also they show their training costs publicly. Can't say for sure since I have not used it personally, but I think they'll for sure outlive the western SOTAs.
          • LogicFailsMe 2 hours ago
            OpenAI apparently runs a profitable inference business with 40% gross margin, but their advertising budget is nutso and their real costs are pretraining and research. I suspect Deepseek's comp is not predicated on capturing the lightcone of all future value, some googling insinuates their top pay is $212K US which would support that suspicion. Compare and contrast with the $1.35M and up at OpenAI.
        • amrit3128 3 hours ago
          Ah yes, I'm sure Trovalds and Stallman are harvesting my data through free software, aren't they? This argument is used by boomers who were fed cold war era propoganda that surely everybody is selfish, and you're always at fault.
          • kipchak 1 hour ago
            Think they're talking about things that are free as in beer but not free as in freedom, not FOSS
      • halJordan 4 hours ago
        It's the size of a big vm. There's nothing wrong with reserving that much working space for one item.
    • cynicalsecurity 3 hours ago
      I don't understand the desire to run own AI models for programming locally. No laptop is ever going to be as powerful and energy efficient to run anything close to OpenAI, Anthropic or Google models. A model you can run on a loptop is simply not going to work as well as it's needed for programming. Small models for linguistic work fine, but anything more sophisticated simply won't provide enough resources or power. Or models would need to be significantly dumbed down - then why use them at all? So far the idea of carrying a "thin" or "thin"-like device looks more reasonable to me, while running AI on your own server.
      • linguae 2 hours ago
        I’m quite optimistic about the long-term future of local LLMs for privacy and cost control reasons. An LLM running on my own hardware, even if it’s not a laptop but a home server, is one where I don’t need to worry about token limits, token fees, privacy, and “rug-pulling” from the vendor.

        In the short term, the big challenge is being able to afford hardware that can run a ~30B model. Last month I got to experiment with LLMs on a NVIDIA RTX 6000 Ada Generation as a visiting researcher during my summer break. I see the power of local LLMs for agentic coding; they’re no Claude, but they are quite useful. I wish I had gotten into local LLMs before hardware has gotten prohibitively expensive and in some cases unavailable; Apple discontinued certain Mac Minis and Mac Studios with high amounts of RAM due to the RAM shortage.

        Hopefully high RAM prices don’t become a new normal, though the next year or two doesn’t look good.

      • OtherShrezzing 3 hours ago
        > A model you can run on a loptop is simply not going to work as well as it's needed for programming

        The models you can run on a high-spec laptop today are approximately where frontier models were 12-18mo ago (albeit at a lower tok/s rate). If you scan back through hn comments from that era, you’ll find plenty of people saying “this is powerful enough to massively increase my productivity”.

        • anon373839 2 hours ago
          > albeit at a lower tok/s rate

          Not always! I get 80-100 tok/s from Qwen 3.6 35B-A3B on a MacBook Pro thanks to MTP. With long contexts that dips to around 50-60. However, prefill is much slower than API models. So it becomes really, really, really critical to not have cache misses.

      • ComputerPerson 3 hours ago
        I've never done it but would be interested because it cuts out the burden of worrying about costs. Maybe I'm mistaken on energy cost here. There's a constant raincloud that follows me around regarding limits, and it would be nice to shake that.

        I've been able to accomplish incredible feats (for myself) since GPT-4, so model intelligence is secondary.

      • brandon272 1 hour ago
        > I don't understand the desire to run own AI models for programming locally.

        Privacy. Security. Not bulk uploading your trade secrets and intellectual property to Sam and Dario’s servers.

      • lluisantoni 2 hours ago
        For some companies there might be a need to run them locally. For instance, Apple decided to run LLMs on the phone locally. I guess it depends on how important latency and privacy are. Perhaps Meta is looking at how much interest for those local models is there.
      • flaburgan 2 hours ago
        Yet.
  • androiddrew 18 minutes ago
    I'd really like to see a 45B-ish dense model ready for a dual GPU setup. Something with a little more intelligence while still within the range of some higher end local setups.
    • tgtweak 13 minutes ago
      There is definitely an under-served target memory size of 48GB - almost everything aims for: 12, 16, 24, 32, 64, ...) But most dual-gpu setups, 3090/4090 (and some mac configs afaik) have 48GB, and most 64GB systems would do well with the extra 16gb of overhead saved. 48GB is also moderately common in PC memory configurations since 24gb DIMMs are a thing.
  • OsamaJaber 2 hours ago
    The comparison set is Gemma4-31B and Qwen3.6-27B, not the current Qwen

    Fair on size, but the headline numbers are against a model a generation back

    • NorwegianDude 16 minutes ago
      That is the most recent Qwen and Google models, there is no newer version, yet. Qwen3.8 27B might come in a couple of days tho, if it's launched alongside the large one when the Qwen3.8 countdown reaches zero.
    • Zambyte 1 hour ago
      What more recent open weight Qwen release is there?
  • Gecko4072 4 hours ago
    What I think would be perfect is a model that could run on a single DGX spark and be competitive with DSV4 Flash 731. Flash is already a game changer. Hopefully meta plans on this, like the old 70b. V4 flash is smart enough for any use but slightly too big. 27b-30b isn’t intelligent enough.
    • cmrdporcupine 4 hours ago
      This model I think will be too slow for that on Spark, even at 4 bit quant.

      It's a dense model, not MoE like e.g. Qwen 35b or Gemma 4 26B A4B. On a Spark it will be memory bandwidth limited

      I haven't tried yet (working on it) but back of the napkin estimate puts it at around 15tok/s even after converting to NVFP4. Prefill would be much higher though. That 15tok/sec is pretty typical for dense models of this size:

      NVFP4 Q/K/V/O and MLP projections: ~13 GB/token

      BF16 attention gates: ~3 GB/token

      BF16 LM head: ~2.5 GB/token

      Total: ~18.9 GB/token

      At 273 GB/s, that gives a bandwidth-only ceiling of about 14.5 tok/s; actual performance would be lower.

      • rao-v 3 hours ago
        Native dflash support on day 1 helps a lot! High quality speculative decoding speeds up a lot of agentic work.
        • cmrdporcupine 1 hour ago
          You're right. I'm getting ~33tok/sec w/ dflash on it, even bursts up to 60tok/sec, using my personal home-built-for-Spark inference engine (not vLLM or llama.cpp based)

          That's pretty respectable.

          Still working on optimizing and cleaning up before I push it.

    • 127 4 hours ago
      DSV4 Flash 0731 already runs on RTX 4090 24GB + 128GB system RAM at a usable tok/s and quantization.
      • Gecko4072 4 hours ago
        You personally? Just curious. Context window is also a factor and ram isn’t really cheap. Sparks are assembled units which I like.
        • dannyw 3 hours ago
          For the same price as a DGX Spark here (A$8499) I can buy roughly 544GB of DDR5-5200MHz from retail; which on a quad channel platform would deliver ~160gb/s real world; and ~320gb/s with octa channels (Xeon, Threadripper Pro).

          If you can afford it or somehow find a used unit, you can go Epyc for 12 channels.

          8/12 channel DDR5 will beat DGX Spark in inference/decode even without a GPU of any kind, as it’s memory bandwidth bound, and the Spark tops out at ~240gb/s real world.

          With some optimisation and maths, it’s entirely plausible to ach

          You are paying an extraordinary amount of money for the convenience of a super small unit, with still mediocre software support, but at least a community. Expect to be crawling through forum posts regularly, as SM121/Spark has many quirks and ecosystem issues still.

          Please don’t pay another 70-80% gross margins on top of already inflated DRAM prices unless you need. The Spark IS really nice if you want to test out ConnectX or if you really need something small and compact and quiet.

          Also consider: used Adas or even Ampere NVIDIA workstation GPUs can come with a lot of VRAM and be “reasonable”, with CUDA.

          • danielEM 1 hour ago
            Been investigating these multichannel AMD based platforms last year and seem like none of them can in real scenarios utilize anywhere close to their theoretical bandwidth.
  • soupspaces 1 minute ago
    what's the catch?
  • mirekrusin 2 hours ago
    Great to see Meta back, looks like really strong, local model, can't wait for llama.cpp support.
    • jakswa 1 hour ago
      some support already merged, and I verified in a local build that it runs (cannot get MTP params working tho, about ~40 tok/s on my beefy 800GB/s 7900XT w/ 20GB VRAM). https://github.com/ggml-org/llama.cpp/pull/26841
      • bwfan123 33 minutes ago
        Cant wait to kick the tires. I am running qwen3.5-coder:35b on my laptop, and while thats nice, such alternatives are welcome.
  • maxignol 4 hours ago
    Optimizing speed is really the way to go. Yet 24GB is not what everyone can afford. Maybe we could take some of those 56tk/s and transfer into some free RAM space using MoE loading ? I'd be glad with a less than 10GB and more than 6tk/s model.
    • lisplist 4 hours ago
      Unfortunately this is just the entry price for LLMs. With the exception of the Qwen 27B models, I personally haven’t found a ton of use cases for models less than 200B. With the right setup, fine tuning, etc, you can make small models do cool things, but hard to please everyone given the insane hardware costs at the moment and the comparably cheap API costs.
      • dannyw 1 hour ago
        Small models are still great for lots of “simple intelligence” use cases, like annotating or summarising files and media; or even just basic chat when given web search tools.

        My local NAS is private and I’m not going to send it off to APIs for captioning or metadata; but even Qwen3VL 8B does an excellent job at this, despite being quite old.

        They are also really excellent for fine tuning. Unsloth and Tinker (from Mira’s TML) are great places to start.

        If your use case is narrower than “coding agent for everything”, you can probably match frontier performances on that narrow domain with ~30b and exceed it with ~100b+.

      • dist-epoch 3 hours ago
        Gemma4-E4B (4B params) works pretty well as a local wiki, or when you don't have connectivity.
        • dannyw 1 hour ago
          Nitpick: Gemma4-E4B is actually a 8 billion param model, but only 4.5B params worth of memory bandwidth needed per decode.
    • Manfrednotfunny 4 hours ago
      I don't thinnk just MoE will solve it. If you hit constantly different expert layers, you can't outsource layers efficently and have to swap it in.

      MoE will be faster because it will read less memory for sure, you still have to have it though.

  • spaqin 1 hour ago
    That's a bit amusing - not that I have the hardware to run it, but officially it's not available in Hong Kong. Not that getting it would be much of a problem with a help of a VPN either, but I'll assume mainland China is also restricted. Certainly not a competition for Chinese open weight models... in China.
  • vibe42 3 hours ago
    Meta released their own 4-bit quant of this model for devices with 24GB VRAM.

    That's a modern gaming laptop; cheapest I see in the US with 24GB is $3.5k.

    Should be quite a bit faster than the new M5 MacBook Pro, and you can run Linux on it!

  • richardfey 4 hours ago
    Looking forward to giving this a try with llama.cpp. I’m watching the open-weights competition with high expectations.
  • tosh 4 hours ago
    good to see new open weights releases from meta
    • jauntywundrkind 4 hours ago
      good looking showing too, which is excellent.
    • InfiniteLoup 4 hours ago
      The least they could do, after ruthlessly bombarding my employer's servers with requests, ignoring the robots.txt, scraping everything, and incurring significant Google Maps costs for us in the process.
  • jakswa 1 hour ago
    Another candidate for the 7900XT (20GB VRAM) I got sitting around. I pulled latest llama.cpp (targeting vulkan during build) after seeing a muse PR merged a few hours ago, and unsloth/Muse-Glimmer-30B-GGUF:UD-Q4_K_XL runs on my 7900XT barely (and with no MTP). Sits at 19GB VRAM w/ 4 parallel 113k context slots, all layers on GPU, and at 700 tok/s prompt, and ~36 tok/s generation.

    Waiting on Q3 to download to check speed + do my usual anecdotes. I generate beefy code snippets and poems, and also ingest my HOA declaration and answer nuanced questions.

    edit: i should've prefaced this somewhere with: This card ballparks at 800GB/s IO, which I can't seem to find easily on the market anymore. Kinda the ideal card for this model, if I just had a _little_ more VRAM (XTX is 24GB).

    edit2: not mtp, this is dflash model (param in child comment). I'm up to ~60 tok/s generation and sitting at 19GB VRAM (i added --no-mmproj (makes it text-only i believe) because I'm used to speculative decoding wanting more VRAM and I'm already close to the limit :sweat_smile:)

    • jakswa 1 hour ago
      Q3 results: unsloth/Muse-Glimmer-30B-GGUF:UD-Q3_K_XL gets down to 15.6GB VRAM and full context (131k) on the 4 parallel slots. Prompt/generation speeds about the same. Overall feeling like a nicer-fitting Qwen 3.6 27B, but want to test out MTP generation speeds once I can.

      edit: My favorite bit of reasoning I saw go by in my "generate me a beautiful code snippet" anecdote: 'Could give a snippet of beautiful code: the "hello world" in brainfuck? No.'

      edit2: my first dflash speculative model! no mtp. I'm up to ~60 tok/s on empty context with `--spec-type draft-dflash`

  • treksis 9 minutes ago
    thank you zuck.
  • Havoc 4 hours ago
    The favourable comparisons to Gemma 4 and qwen3.6 look promising!
    • cmrdporcupine 4 hours ago
      Those two offer MoE variants, this doesn't seem to.

      Dense model makes it dog slow on anything without HBM. Max 15tok/sec on decode on DDR5 systems like a Spark or a Strix Halo -- and that's at 4 bit quant.

      • EddieRingle 3 hours ago
        Dense models run at a very usable speed (Qwen 3.6 was running at ~50t/s last I looked) on my dual 7900 XTX desktop. (And before anyone brings it up, I did not buy them for this purpose, so the up-front cost is irrelevant in my case.)
      • Havoc 2 hours ago
        The benchmark comparison is against the dense variants not MoE
      • petu 3 hours ago
        3090/4090 probably would do 40 t/s, for 5090 75 t/s is shown in the blog.
  • bentt 3 hours ago
    Meta seems like the one American bigtech that would distill the the other American frontier models. My enemy’s enemy is my friend?
    • Maxious 1 hour ago
      > Some have tried to frame distillation as harmful, but I think it is important to protect the principle that you can learn from anything you can observe.

      - Mark Zuckerberg

      https://www.meta.com/thefutureisforeveryone

    • dev_daftly 1 hour ago
      You think the company buying up all the books, cutting off the bindings, and feeding them through a scanner isn't also distilling other models?
    • grim_io 3 hours ago
      They do distill, their own bigger Muse model.
  • hndhyc0bdt 1 hour ago
    Refreshingly practical
  • jckahn 1 hour ago
    Where is the pelican??
    • jakswa 28 minutes ago
      I'm listening to pelican sounds on youtube while I wait for Simon.
  • wyzer 1 hour ago
    How are you handling the tradeoff between quantization for device fit and accuracy loss on tool calling? That's where local agents typically break down in production.
  • brumbelow 8 minutes ago
    and now the recent Meta model 'security issue' begins to make sense
  • mytailorisrich 32 minutes ago
    Random question: Would you be able to run this model on a Macbook Air M5 (latest)?
  • gunalx 4 hours ago
    Meta did not abandon opensource. I would love to see a smaller distill, or a moe of this size but the benchmarks seems competetive as long as it isnt benchmaxed witch i would not be suprosed if it is.
    • ignoramous 3 hours ago
      > Meta did not abandon opensource

      Open weights*

      I don't think outside of the Big 3 (Ant, OAI, GDM), given the strong competition from China, any other Lab has a chance at capturing the coding market if they aren't open weights (save for xAI whose latest Grok looks every bit good & will probably rely on Cursor for distribution instead of going open weights). There's literally no other selling point, as the capabilities have mostly converged by now among the chasing pack.

      • dannyw 1 hour ago
        Don’t sleep on NVIDIA and Nemotron.

        It’s not completely open source, but they actually release their pretraining and post-training datasets with some redactions for (cough) pirated content.

        They also have very good code and playbooks for actually doing a fine-tune, CPT, etc.

        Even if you’re not tuning a Nemotron model, its mixes are very excellent for your replay data slice; or general experiments. Way better curation and quality than Dolma, etc; or other large huggingface data mixes I tested.

        • nickludlam 1 hour ago
          Yes, I second Nemotron. I'm using Ultra remotely and Super locally, and I find them very useful for RAG-like problems. I wouldn't really use them for coding.
      • ComputerPerson 3 hours ago
        There was a good discussion yesterday on the DeepSeek Flash release thread about this.

        There's a large market, very large, who want the best regardless of what it costs. Probably a large enough market to keep that domain of research afloat (as opposed to shifting research manpower to cost cutting).

        The reasoning is just that the marginal cost of AI is very secondary to fixed costs of the businesses themselves; it's not an excuse to sacrifice performance.

      • HardCodedBias 54 minutes ago
        GDM -- Ok, I'll bite. Why are you including them?
        • dannyw 48 minutes ago
          Some labs go through bad patches, GDM is definitely in one right now and the recent departures are not reassuring, but I think it's too early and dismissive to count them out of the race so far. They just need one good frontier release for everyone to go "GDM is back!"

          Claude models weren't really good or noteworthy until the 3 series anyway.

          • HardCodedBias 31 minutes ago
            I think the record is quite clear, GDM was never in the race.

            All of the Gemini models have been considerably behind the capabilities frontier. The only exception was 3.0 which seemed quite good, but had latent issues and we were all measuring with the incorrect metric, agentic where it's latent issues were very pronounced.

            GDM+Google may have created an exceptionally efficient LLM for serving search. This is likely a great accomplishment (or maybe Google is burning money at a rate unheard of before). But Frontier capability: they have never been in the race.

            This is sad, since they had everything necessary to be on or beyond the frontier.

  • solarkraft 4 hours ago
    Wow, Meta is back (at least for now)!

    I like this class of model. Multi-token prediction makes it viable to run dense models at not-too-far-off speeds as MoE models with much better intelligence.

    The submission’s title (open weights 30B local coding model) is luckily wrong: This is meant to be a general agentic model.

    It even comes pre-quantized and with a MTP/drafter model. Looking good!

    Let’s hope they aren’t dishonest with the benchmarks this time …

    • akazantsev 1 hour ago
      > The submission’s title (open weights 30B local coding model) is luckily wrong: This is meant to be a general agentic model.

      https://xcancel.com/alexandr_wang/status/2086756152034066792

      It's correct. See the OpenCode demo. Generic models are good enough for coding without necessarily being designed specifically for coding.

    • bwfan123 19 minutes ago
      > It even comes pre-quantized and with a MTP/drafter model

      Glad to see the extra engineering effort that went into creating this local model and making it run well on a consumer device. I use qwen3.5-coder, and am waiting to kick the tires on this one. I hate to say this, but kudos to Meta ! I hope apple and others follow suit and create similar local models for other use cases like audio, images and video that can run on a laptop.

  • bronxbomber92 4 hours ago
    I wish they would release the quantized versions in a safetensor format. Many frameworks can't load PTE and GGUF.
  • zmmmmm 4 hours ago
    Meta knows how to win back developer's hearts .... let's see if they have the goods
    • xandrius 4 hours ago
      If there is anything meta can do to regain hearts other than owning up their evil deeds, radically change their business model and paying up for taxes and damages, then the world is truly fucked and corporations will continue to win.
  • HardCodedBias 1 hour ago
    LOL the mogging of GDM is hilarious.

    I don't know why MSL released this, but it is very nice that they did.

  • nutjob2 4 hours ago
    The more open weight models get released the greater the market for personal and small business oriented hardware to run these models. This will drive lower cost hardware, which has stagnated in recent years due to most software not needing the performance and capacity.
    • grim_io 3 hours ago
      Higher demand for 5090's did not make them cheaper, because Nvidia got much higher margin products to focus on.
    • cmrdporcupine 3 hours ago
      The opposite happening because foundries are full to capacity making higher margin stuff.
  • sgt 3 hours ago
    Can I run this on my RTX 5090?
    • skohan 2 hours ago
      Yes they have quants for 32GB and 20GB use-cases (including mmproj and kv cache + context)
  • eugene3306 2 hours ago
    will it run on 2x 5060Ti with 16GB each?
    • leansensei 47 minutes ago
      It does, beautifully. Now let's wait for an NVFP4 GGUF!
    • skohan 2 hours ago
      It should - the kquant-dynamic variant is targeted towards 32GB. Downloading it now to give it a try.
  • TommyLe999 34 minutes ago
    [dead]
  • ed 1 hour ago
    [dead]
  • TommyLe999 34 minutes ago
    [dead]
  • korykaai 2 hours ago
    [flagged]
  • jkwang 4 hours ago
    [flagged]
  • moron4hire 2 hours ago
    "Meta Muse" immediately made me think of Metamucil.

    Product teams really need to hire at least one or two people with a 12-year-old's sense is humor. They need to winnow all the potential stupid jokes out of their product namings.

  • dhchun1203 1 hour ago
    Three of these landed in the same week. Mistral's Shieldstral is a 3B safety classifier that matches models 7x its size, and Google shipped Gemma Translator which runs entirely offline. Different problems, same shape. Small open weights, local, no API call.
  • petcat 3 hours ago
    As an industry, I wish we would stop calling these things "open weight" because it is too easy to confuse with actual "open source", which they are not.

    Photoshop source code+ OSI license = open source

    Photoshop binary you can run on your own computer = open weight

    Photoshop SaaS web app = closed, proprietary (Opus, GPT, etc.)

    "Open weight" models are still just binary blobs that are completely inscrutable. It's like bringing home a dog from the rescue and just hoping that it doesn't have a tendency to bite kids in the face. You just can't know. The only thing that you can do is try to add more training (fine tuning) telling it not to bite kids.

    I don't think the FOSS community has ever accepted this, but somehow we're feeling like it is okay now.

    • microtonal 3 hours ago
      Photoshop source code+ OSI license = open source

      Photoshop binary you can run on your own computer = open weight

      I don't think this is a correct analogy. You are not allowed to distribute modified versions of the Photoshop binary. Most open weight model licenses allow you to make and distribute your own finetunes, etc.

    • QuadmasterXLII 3 hours ago
      Given an open weights model trained to sometimes bite kids, we can’t train it to not bite kids, even though billions of dollars of research have been thrown at this open problem.

      Given an open weights model trained to never bite kids, you can get it to bite kids with 10 prompts and a linear projection, the known simple algorithm doesn’t even need a backwards pass.

      yay asymmetry!

      • kzrdude 3 hours ago
        Any pointers to more info about that? Sounds interesting.
    • craigmart 3 hours ago
      I believe that comparing LLMs with traditional deterministic software is fundamentally misleading. It is extremely difficult to truly interpret what LLMs do internally, and as of now, nobody fully understands it. Even if you trained the LLM yourself, there is no source code you can simply read and learn from.

      Sure, having information about how these models were trained is helpful for reproducibility, but it is basically impossible for anyone without substantial capital and access to the same (likely copyrighted) data to reproduce the model. For normal users, owning the model weights essentially means owning 100% of the model, you can inspect and study the weights in much the same way as the lab that produced the model can, you can modify the weights, and you can use and distribute them if the license allows you to

    • piker 3 hours ago
      It is useful to indicate you can run the weights on your own hardware. That’s categorically different from most other commercial offerings. It’s as if your adobe example ignores the reality that would exist had photoshop been invented in 2019: cloud only.
    • monster_truck 3 hours ago
      This analogy is terrible and seems to be extremely misinformed about how rescues evaluate dogs before they are put up for adoption
      • petcat 3 hours ago
        I am extremely well aware of how rescues evaluate dogs. And I'm also fully aware that they do not know the full history of the dog. They go through a limited set of testing and interrogation to evaluate the safety of the dog. That's it.