AgentStack
SKILL verified MIT Self-run

Neural Network Design

skill-ericwang915-data-scientist-skills-neural-network-design · by ericwang915

Design state-of-the-art neural network architectures: TabNet/TabPFN for tabular, Vision Transformers/ConvNeXt V2 for images, Mamba/RWKV state space models for sequences, Mixture of Experts, diffusion models, and multimodal architectures. Includes Flash Attention, LoRA, quantization, and modern training recipes. Use when designing a neural network for any modality.

No reviews yet
0 installs
11 views
0.0% view→install

Install

$ agentstack add skill-ericwang915-data-scientist-skills-neural-network-design

✓ scanned · ✓ verified — works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

Are you the author of Neural Network Design? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Neural Network Design

Purpose

Design production-grade neural architectures using the latest SOTA techniques from top ML labs.

How It Works

Architecture Selection (2024-2026 SOTA)

| Problem | SOTA Architecture | Why | |---------|-------------------|-----| | Tabular | TabPFN, XGBoost, TabNet, FT-Transformer | TabPFN: zero-shot, no tuning; FT-Transformer: attention on features | | Images | ConvNeXt V2, DINOv2, EVA-02, SigLIP | DINOv2: best self-supervised; SigLIP: multimodal alignment | | Text | Llama 3, Mistral, Qwen 2.5, Gemma 2 | Open-weight LLMs for fine-tuning | | Sequences | Mamba-2, RWKV-6, xLSTM, Hyena | State space models: O(n) vs O(n²) attention | | Audio | Whisper v3, Encodec, VALL-E | Whisper: robust ASR; Encodec: neural audio codec | | Video | VideoMAE v2, InternVideo2, Sora-like diffusion | Temporal understanding + generation | | Multimodal | LLaVA-NeXT, GPT-4V architecture, Gemini pattern | Vision encoder + LLM backbone + projection | | Generation | Stable Diffusion 3, FLUX, Consistency Models | Flow matching > DDPM; Consistency: single-step | | Graph | GPS (General Powerful Scalable), GraphMAE | Positional encodings + attention on graphs |

Modern Building Blocks

Attention Variants
Standard Self-Attention:   O(n²) memory — vanilla transformer
Flash Attention v2:        O(n) memory — fused CUDA kernel, 2-3x faster
Grouped Query Attention:   Share KV heads across Q heads (Llama 2+)
Sliding Window Attention:  Local context + global tokens (Mistral)
Linear Attention:          O(n) via kernel approximation (Performer, RWKV)
Ring Attention:            Distribute context across GPUs for >1M tokens
Normalization & Activation
  • RMSNorm > LayerNorm (simpler, equally effective — Llama, Mistral)
  • SwiGLU > GELU > ReLU (gated activation, +1-3% on LLMs)
  • Rotary Position Embedding (RoPE): Relative positions via rotation matrices
Efficient Architectures
  • Mixture of Experts (MoE): Route tokens to specialized sub-networks (Mixtral, Switch Transformer)
  • 8 experts, top-2 routing → 8x parameters at 2x compute
  • State Space Models: Mamba — selective scan, hardware-aware, O(n)
  • Better than Transformers on long sequences (>8K tokens)
  • Hybrid architectures: Mamba + Attention layers (Jamba, Zamba)

Tabular Deep Learning (Often Overlooked)

| Model | Type | When | |-------|------|------| | TabPFN | In-context learning | 10K rows, categorical-heavy | | CatBoost | Ordered boosting | High cardinality categoricals |

Reality check: On most Kaggle tabular benchmarks, well-tuned XGBoost still beats deep learning. Use DL for tabular only when: multi-modal inputs, pre-training benefits, or entity embeddings add value.

Training Recipe (Modern Best Practices)

# Production Training Config (2025 Standard)
config = {
    "optimizer": "AdamW",           # or Adafactor for memory savings
    "lr": 3e-4,                     # base LR; scale with batch size
    "lr_schedule": "cosine_with_warmup",
    "warmup_ratio": 0.06,           # 6% of total steps
    "weight_decay": 0.1,            # decouple from LR
    "grad_clip": 1.0,               # gradient clipping norm
    "precision": "bf16",            # bfloat16 > fp16 (no loss scaling needed)
    "compile": True,                # torch.compile for 20-30% speedup
    "gradient_checkpointing": True, # trade compute for memory
    "fsdp": True,                   # Fully Sharded Data Parallel for multi-GPU
    "activation": "swiglu",
    "norm": "rmsnorm",
    "pos_encoding": "rope",
}

Parameter-Efficient Fine-Tuning (PEFT)

| Method | Trainable Params | Memory | When | |--------|-----------------|--------|------| | Full fine-tuning | 100% | High | Lots of data, large GPU | | LoRA | 0.1-1% | Low | Standard PEFT, most cases | | QLoRA | 0.1-1% | Very low | Fine-tune 70B on single GPU | | DoRA | 0.1-1% | Low | Better than LoRA on some tasks | | Prefix tuning | <0.1% | Very low | Prompt-like, simple tasks | | Adapters | 1-5% | Low | Multi-task, switching |

Usage Examples

"Design a multimodal model that takes product images + descriptions
and predicts category — I have 50K labeled examples"
"What's the best architecture for processing 100K-token documents?
Transformer attention is too expensive."

Output Format

  • Architecture: Layer-by-layer spec with modern components
  • Training Recipe: Optimizer, scheduler, precision, parallelism
  • Hardware Estimate: GPU memory, training time, cost
  • PyTorch Code: Complete model + training loop with torch.compile
  • Alternatives: Trade-off analysis (speed vs. quality vs. cost)

Key References

  • Gu & Dao (2023) — Mamba: Linear-Time Sequence Modeling with Selective State Spaces
  • Dao (2023) — FlashAttention-2: Faster Attention with Better Parallelism
  • Hollmann et al. (2023) — TabPFN: A Transformer That Solves Small Tabular Classification Problems in a Second
  • Fedus et al. (2022) — Switch Transformers: Scaling to Trillion Parameter Models

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet — be the first.

Versions

  • v0.1.0 Imported from the upstream source.