— No reviews yet
0 installs
11 views
0.0% view→install
Install
$ agentstack add skill-cxcscmu-skilllearnbench-simpo-loss-function ✓ scanned · ✓ verified — works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Are you the author of Simpo Loss Function? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claimAbout
SimPO Loss Function Implementation
Overview
SimPO (Simple Preference Optimization) implements a preference optimization objective that uses length-normalized average log probability as an implicit reward, with a target reward margin component.
Key Formula
L_SimPO(πθ) = -E_(x,yw,yl)~D log σ(β/|yw| log πθ(yw|x) - β/|yl| log πθ(yl|x) - γ)
Components
1. Length-Normalized Reward
- Formula:
r_SimPO(x, y) = β/|y| * log πθ(y|x) - Purpose: Average log probability per token, prevents length bias
- Why: Aligns training with generation metric (which uses average log likelihood for beam search)
2. Bradley-Terry Objective
- Formula:
p(yw ≻ yl | x) = σ(r(x, yw) - r(x, yl) - γ) - Purpose: Probabilistic ranking between winning and losing responses
- Function: σ is sigmoid function
3. Target Reward Margin (γ)
- Purpose: Ensure reward difference exceeds a target threshold
- Effect: Improves generalization by enforcing margin between classes
- Typical range: 0.5 to 1.5
Implementation Details
Computing Log Probabilities
# log_probs shape: (batch_size, seq_len)
# Sum across sequence dimension to get total log probability
log_prob_sum = log_probs.sum(dim=1) # (batch_size,)
# Divide by sequence length for normalization
seq_lengths = (input_ids != pad_token_id).sum(dim=1) # (batch_size,)
avg_log_prob = log_prob_sum / seq_lengths.float() # (batch_size,)
Computing Reward Differences
# Batch structure: pairs of (winning, losing) responses
batch_size = avg_log_probs.shape[0]
winning_rewards = avg_log_probs[:batch_size//2]
losing_rewards = avg_log_probs[batch_size//2:]
# Reward difference with margin
reward_diff = beta * winning_rewards - beta * losing_rewards - gamma
Computing Loss
# Bradley-Terry with sigmoid
import torch.nn.functional as F
sigmoid_term = torch.sigmoid(reward_diff)
loss = -torch.log(sigmoid_term).mean()
Common Pitfalls
- Not using length normalization: Creates bias toward longer sequences
- Wrong batch structure: Ensure paired winning/losing responses
- Missing average in log probabilities: Use sum/length, not just sum
- Gradient flow: Ensure no detach() breaks gradients to model
Hyperparameters
- β: Temperature parameter, typically 2.0-2.5
- γ: Target margin, typically 0.3-1.6, depends on setting
- learning_rate: Usually small, 1e-6 to 5e-7
References
- SimPO Paper: Section 2.3 "The SimPO Objective"
- Length normalization: Equation (3) in paper
- Gradient analysis: Appendix F
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: cxcscmu
- Source: cxcscmu/SkillLearnBench
- License: MIT
- Homepage: https://cxcscmu.github.io/SkillLearnBench
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet — be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.