Let me investigate the tokenizer behavior to find the root cause. Now let me check the actual token IDs for the letters "A", "B", "C", "D": Let me investigate the tokenizer behavior directly. Let me verify the exact tokenizer behavior with single letters using a direct tokenization check. ## Found it! 🐛 The bug is in how the ByteLevel BPE tokenizer + logit slicing interact. Let me verify my hypothesis by checking the actual vocab tokens: Now I have the critical data. Let me verify the byte-level encoding behavior by checking the ByteLevel-specific token mappings. The bug is **not** in the MLP logic or the data processing — it’s in how the LLM’s output is interpreted. **The root cause:** The code only examines the logits (raw prediction scores) for the four single-letter tokens `'A'`, `'B'`, `'C'`, `'D'` at the **very first token** of the assistant’s response. But the model is not forced to start its reply with one of those letters. In practice, SmolLM2 often begins its response with a **space, newline, or special token** (e.g., after `<|im_start|>assistant\n`). When that happens, the logits for the four letters are still computed, but their relative ordering is dominated by the **prior probability** of the tokens (alphabetical bias) rather than by the actual performance scores in the prompt. Because `'A'` is the first option and has a slightly higher prior than `'B'`, `'C'`, `'D'`, it wins the argmax **every time**, regardless of the score values. **Why it selects RELU (Strategy ALPHA) 100% of the time:** - The first token the model *wants* to output is rarely `'A'` (it's often a space or newline). - Among the four letter tokens, `'A'` consistently has the highest logit due to the order in which the prompt lists the options and the model’s internal biases. - The scores in the prompt (which should influence the decision) are **ignored** because the model’s first‑token prediction does not rely on them — it has already “decided” to output something else first. **The fix:** Instead of taking the argmax of logits for a fixed set of tokens, **generate the actual next token** (or a short sequence) and check if it matches one of the letters. If it doesn’t, continue generating until a letter appears, or use soft‑max sampling over the four letters *after* accounting for the model’s true first‑token distribution. Example corrected approach: ```python with torch.no_grad(): outputs = llm_model(**inputs) next_token_logits = outputs.logits[0, -1, :] # Sample from the full distribution, not just the four letters next_token_id = torch.multinomial(torch.softmax(next_token_logits, dim=-1), 1).item() # If it's not a letter, force sampling from the letter subset or re‑run? # Better: use generate() with prefix_allowed_tokens_fn to restrict output to letters ``` This would let the model actually emit the letter it chooses, and the selection would depend on the performance scores instead of a fixed alphabetical prior.