MediumPhase 7 — LLM training, RAG, agents, and evaluation

Week 73 — Alignment

LLM training, RAG, agents, and evaluation · Read / diagram

Medium

Phase 7 — LLM training, RAG, agents, and evaluation

Overview

RLHF, DPO, ORPO, GRPO

Mode this week: read_diagram. Match the work to that mode (operating model): implement ships running code; read/diagram ships a doc; deploy/benchmark ships a measured run. Personal dates, checkboxes, and the RPG layer stay in a learner journal.

Core deliverable

Preference eval plan

Optional depth

Add a production-shaped benchmark, cost or safety analysis, and an architecture trade-off note.

Week loop (from the cookiecutter journal)

Forge theme: KMP, Aho–Corasick, trie retrieval, sequence DP, suffix-array survey

Leetcode Darbar (parallel)

This week's band (70-81): Strings, tries, graphs, sequence DP. Production connection: Tokenization, retrieval, agents, guardrails.

Standard: one aligned problem. Auror: the five-slot queue on the Darbar page. Tracker stays in the learner journal.

Topics: graphs · string · trie

Agent papers

This week sits on the GenAI agents track. Read the assigned paper and connect it to the focus above.

Evidence contract

Every core week records all four (cookiecutter journal/weeks/week_073.md):

  1. Code / implementation (or the design artifact on a writing week)
  2. Benchmark / result
  3. Design doc / technical explanation
  4. Retrospective / learning note (Sailboat: destination, wind, anchor, rocks, heading)

On this site

Official Tensor-to-Tenant


Week 73 of Tensor-to-Tenant · Previous week · Next week.