AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Simpo Loss Function

skill-cxcscmu-skilllearnbench-simpo-loss-function · by cxcscmu

Implement SimPO loss with length-normalized rewards and target margin.

— No reviews yet
0 installs
34 views
0.0% view→install

Install

$ agentstack add skill-cxcscmu-skilllearnbench-simpo-loss-function

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • ✓ Prompt-injection patterns
  • ✓ Secret / credential exfiltration
  • ✓ Dangerous shell & filesystem operations
  • ✓ Untrusted network calls
  • ✓ Known-malicious package signatures

What it can access

  • ✓ Network access No
  • ✓ Filesystem access No
  • ✓ Shell / process execution No
  • ✓ Environment & secrets No
  • ✓ Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-cxcscmu-skilllearnbench-simpo-loss-function)

Reliability & compatibility

✓ Security review passed
0 installs to date
— no reviews yet
● 3mo ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Simpo Loss Function? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

SimPO Loss Function Implementation

Overview

SimPO (Simple Preference Optimization) implements a preference optimization objective that uses length-normalized average log probability as an implicit reward, with a target reward margin component.

Key Formula

L_SimPO(πθ) = -E_(x,yw,yl)~D log σ(β/|yw| log πθ(yw|x) - β/|yl| log πθ(yl|x) - γ)

Components

1. Length-Normalized Reward

  • Formula: r_SimPO(x, y) = β/|y| * log πθ(y|x)
  • Purpose: Average log probability per token, prevents length bias
  • Why: Aligns training with generation metric (which uses average log likelihood for beam search)

2. Bradley-Terry Objective

  • Formula: p(yw ≻ yl | x) = σ(r(x, yw) - r(x, yl) - γ)
  • Purpose: Probabilistic ranking between winning and losing responses
  • Function: σ is sigmoid function

3. Target Reward Margin (γ)

  • Purpose: Ensure reward difference exceeds a target threshold
  • Effect: Improves generalization by enforcing margin between classes
  • Typical range: 0.5 to 1.5

Implementation Details

Computing Log Probabilities

# log_probs shape: (batch_size, seq_len)
# Sum across sequence dimension to get total log probability
log_prob_sum = log_probs.sum(dim=1)  # (batch_size,)

# Divide by sequence length for normalization
seq_lengths = (input_ids != pad_token_id).sum(dim=1)  # (batch_size,)
avg_log_prob = log_prob_sum / seq_lengths.float()  # (batch_size,)

Computing Reward Differences

# Batch structure: pairs of (winning, losing) responses
batch_size = avg_log_probs.shape[0]
winning_rewards = avg_log_probs[:batch_size//2]
losing_rewards = avg_log_probs[batch_size//2:]

# Reward difference with margin
reward_diff = beta * winning_rewards - beta * losing_rewards - gamma

Computing Loss

# Bradley-Terry with sigmoid
import torch.nn.functional as F

sigmoid_term = torch.sigmoid(reward_diff)
loss = -torch.log(sigmoid_term).mean()

Common Pitfalls

  1. Not using length normalization: Creates bias toward longer sequences
  2. Wrong batch structure: Ensure paired winning/losing responses
  3. Missing average in log probabilities: Use sum/length, not just sum
  4. Gradient flow: Ensure no detach() breaks gradients to model

Hyperparameters

  • β: Temperature parameter, typically 2.0-2.5
  • γ: Target margin, typically 0.3-1.6, depends on setting
  • learning_rate: Usually small, 1e-6 to 5e-7

References

  • SimPO Paper: Section 2.3 "The SimPO Objective"
  • Length normalization: Equation (3) in paper
  • Gradient analysis: Appendix F

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.