real_faxenoff 1 hour ago

As a regular user of a bunch of specialized micromodels, I'll tell you this: you won't be happy with such a model (and its JEV counterparts) running permanently in the background on your PC's CPU. You need to offload their processing to the NPU. There are many pitfalls along the way, but the result is worth it.

NPU performance will be twice as high, while power consumption will be four times lower. No additional fan noise (if you know what I mean).

I'll wait another month until the first phase of the =battle royale= among models of this kind wraps up, put together a solution for the NPU/iGPU, and post it on HF.

girvo 10 minutes ago

Does anyone know if it is worth fine-tuning one of these decision models on the shape of the questions you want it to work on, vs the more general versions? I'm using Jev pretty successfully at work at the moment, but am curious about what is doable

adenta 4 hours ago

At this point I can't wait for a comedian to release a decision model backed by humans.

Meet Jerry- it's literally a guy named Jerry answering your questions.

mattvr 2 hours ago

Why is everyone calling binary choices `noul`? Does this have some meaning or is it just copying Jev’s API?

  • WASDx 30 minutes ago

    The new OpenAI Decisions API calls it "predicate". Also calling the API "decisions" rather than "system one". Usually I don't like inventing new standards but I hope the OpenAI schema takes over. We don't need this hype terminology.

  • dprkh 14 minutes ago

    Claude generated it and it stuck.

  • tchalla 7 minutes ago

    The same reason why they’re calling this System One thinking. Everyone wants to be seen doing different things and smart ones.

mynti 23 minutes ago

Can someone explain this architecture a bit more in depth? They say the pointer head scores the hidden state at each option against the hidden state of the answer. But the LLM produces hidden states per token, so an option can span multiple tokens, no?

keyle 5 hours ago

Fantastically well written. It's rare for me to be able to understand what the AI gurus are talking about, and this was written by humans for humans.

It can technically be used for a lot of use cases, I'd like people to chime in on ideas on this?

  • nryoo 3 hours ago

    Picking lunch menu..?

SubiculumCode 2 hours ago

Are any of these multimodal yet? I'd love to try asking a model with calibrated probabilities to answer question like, "do these shapes match?". Sure, you can ask a LLM....

  • sauhsoj 2 hours ago

    Strands Decider can take vision in. How does it go with that question?

    • Zopieux 5 minutes ago

      This is not advertised on their page, did you make this up?

      I believe image classification/analysis by deciders (not just OCR, not everything is about text) is still lacking.

      Cloudflare's Clef had fair results on my test, but it's larger and slower. Wondering about Strands.

hrpnk 1 hour ago

clef from cloudflare runs on llama.cpp - being locked-in to strands cli would be a bummer and will slow down adoption.

Since it's a LoRa on Qwen, I assume this is runnable via llama.cpp. Pity that the PEFT/LoRa->GGUF translation is left to the user. Anyone got past:

    $ uv run --with transformers==5.19.0 convert_lora_to_gguf.py ~/Downloads/lora --dry-run --verbose
    [...]
      File "/Users/user/repos/llama.cpp/conversion/base.py", line 630, in map_tensor_name
    raise ValueError(f"Can not map tensor {name!r}")
    ValueError: Can not map tensor 'layers.0.linear_attn.in_proj_a.weight'
soltanov 3 hours ago

Benchmark calibration does not establish reliability on unfamiliar production inputs.

teruakohatu 3 hours ago

Any idea how well this would run on a CPU?

  • gopalv 2 hours ago

    On my M3 mac, it works okay inside a docker container with just CPU.

    { "model": "strands-decider-2B-hobson-v19", "answers": { "is_urgent": { "type": "noul", "noul": 0.8287 } }, "usage": { "input_tokens": 86, "output_tokens": 1 }, "latency_ms": 1732.17 }

    This is how I got it running - https://gist.github.com/2891eb0db9ea92c1a4e860d44f556292

    There's a lot more to be done if we optimize for MLX & let it run on a Mac mini instead of the docker wrapper.

  • avereveard 2 hours ago

    About half a second per decision on six cores

stephantul 2 hours ago

2B being called small is such a sign of the times

davvie 3 hours ago

Looks really nice, I think I could use it on my Mac mini for some smaller automations

yieldcrv 3 hours ago

a strand type game