• Hacker News
  • new|
  • comments|
  • show|
  • ask|
  • jobs|
  • ronef 13 hours

    Asking this with a bias since I work on Nixos.org and Flox.dev - How is the team thinking about the infra layers underneath these models? Any priority or reason to imbed determinism/reproducibility at the bottom of the stack?

  • zurfer 16 hours

    more interesting link: https://arxiv.org/html/2607.09424v2 and > Long-context serving efficiency. Soofi S combines frontier-level capability with the highest measured aggregate long-context decode TPS, and unlike full-attention dense baselines maintains high throughput as context grows. Panel (1(a)) plots Capability Index versus measured aggregate decode TPS/GPU at 40K context and batch 32. The Capability Index averages five benchmark groups, i.e., Code, GSM8K, GPQA-Diamond, English aggregate, and German aggregate, after normalizing each group to the best plotted model. Aggregate decode TPS/GPU is measured with a TP=1, one-B200 vLLM latency-subtraction protocol. Panel (1(b)) shows measured aggregate decode TPS/GPU as a function of input context length under the same batch-32 protocol.

    it's a small win in the small model class

  • karussell 15 hours

    Previous interesting discussion 3 days ago: https://news.ycombinator.com/item?id=48937756

  • myshapeprotocol 15 hours

    Love seeing more emphasis on sovereign open-source models. The shift away from centralized, static credential/identity layers toward self-contained architectures is definitely where the ecosystem needs to head.

    Incipient 14 hours

    But can it, really? It takes a huge amount of time, resources, and knowledge to train a highly capable model. It also takes a huge amount of the same to run them.

    Is it ever going to not be centralised?

  • MSkill1 16 hours

    I'm not seeing how this project is open source exactly. It says license free, but that's just like ChatGPT. I wouldn't call that transparent. Maybe I'm missing something. Google had some difficulty translating the site from German.

    spmurrayzzz 15 hours

    They've open sourced some of the training and inference code: https://github.com/soofi-project

    Their work is based on the nemotron arch (so far).

    zurfer 16 hours

    https://github.com/soofi-project/Soofi-Pretraining https://huggingface.co/Soofi-Project/Soofi-S-Base > "The final model will be released openly under a permissive license, without gated access. We will share access details as soon as it is ready."

    not open weight yet

  • 23ah-qwd 16 hours

    https://www.soofi.info/soofi-s/

    "digitalen Wertschöpfung" (digital value creation)

    Please Jörg Bienert, fuck off. You have never created anything in your life so you do not see or care about the theft. All you do is grin on a photo.

    JSR_FDED 13 hours

    Is there a story behind this?

  • gmerc 15 hours

    This is the Nvidia engagement team coaching various countries (See also Malaysia) how to train a Nemotron and add the benchmark questions to the training set so you get that nice PR splash of pretending it’s the best in its language.

    It’s just nemotron with benchmark juicing and the sovereign smokescreen on good old Nvidia hardware chain remains intact.

    11 hours

  • kvisner 12 hours

    A note on your website, when I hit the EN button in the top right, it presents a pop-up in German, which I can't read, so I can't change it to be in English.

    sajithdilshan 12 hours

    It’s a pop up saying that they use a third party service to translate the website and they collect the activity of the user( which I don’t understand, why would they collect any user data to translate a static website). But it’s typical German like data protection and all the shenanigans.

    Also the UI of the buttons on the popup is terrible both accept and reject buttons looks the same

    burgerone 11 hours

    You should be glad that the popup isn't using any dask patterns to coerce you into choosing an option you wouldn't otherwise agree with