π§ Concepts¶
The ideas behind SLMCode β so the rest of the docs feel inevitable instead of magical. Also: fewer βwhy is it like this?β Slack threads. π
01 Harness β model π§°¶
A model predicts tokens. A harness decides what the model sees, when it acts, how failures recover, and where knowledge sticks.
Frontier tools often hide the harness behind a chat box. SLMCode makes the harness explicit β like leaving the kitchen lights on.
| Layer | Owner | Job |
|---|---|---|
| π§ Routing / board | Go | Plan, schedule, stop/resume |
| π§© Specialists | Prompts + tools | One role, one pack |
| π Critic | Reviewer + disk evidence | Catch fiction |
| πΎ Memory | .slmcode/*.md | Compound lessons |
π€ UnicoLab watercooler
Models are the talent. Harnesses are the stage managers. Never let the talent rearrange the set mid-show.
02 Scoped packs (the turkey rule) π¦¶
Stuffing the whole repo into context is how small models fall asleep mid-sentence.
Each specialist receives a TaskPack:
- slices of PROJECT / CONTEXT / MEMORY / skills
- a few focus files
- one atomic task
- tool allowlists that match the role
flowchart TD
Repo[Whole repo π’] -.->|never wholesale| Model[Model π΄]
Pack[TaskPack π¦] --> Model2[Model π]
Pack --> MD[.slmcode markdown]
Pack --> Files[Focus files]
Pack --> Skill[Matched skills]
Pack --> Task[Atomic task] Bigger models still benefit: less noise, clearer acceptance criteria, cheaper runs. (Your CFOβs favorite sentence.)
03 Plan β split β coordinate π¶
query
β instructions (AGENTS.md / PROJECT.md)
β skills match
β context agent
β explore OR reuse memory
β clarify (interview: ask|auto recommended β Locked PRD)
β scope judge (every task gets concrete acceptance / PRD)
β planner (multipass) β splitter β sanitize (+ auto tester task)
β coordinator advice
β parallel execute (worker smoke + acceptance smoke + static/claims)
β review β correct (β€ max_retries)
β escalate HITL if stuck (timeout β @escalate SLM decides)
β placeholder polish β completeness bar
β finalize tester (real commands required)
β QA gate (install deps + pytest preferred β not compileall alone)
β continue-ask if work remains
β learn β evolve skills β session snapshot
The coordinator doesn't write code. It steers the kanban: promote, reassign, add tasks, note risks. Think air-traffic control, not pilot. βοΈ
04 Explore reuse β»οΈ¶
Deep exploration is expensive (especially on slow local inference).
If CONTEXT is rich, MEMORY/PROJECT exist, and discovery finds relevant paths, SLMCode skips the deep dive and reuses knowledge. Your fans thank you. Your GPU fans thank you louder.
05 Self-critic with evidence π¶
Reviewers can be flaky on SLMs. Heuristics prefer:
- clear
status: done files_changedthat match disk- rename satisfaction when paths already moved
π Disk beats vibes
Always. If the file says hello and the model says goodbye β trust the file.
06 Knowledge flywheel π¦¶
After a run:
- MEMORY.md β lessons / pitfalls
- CONTEXT.md β what we touched
- SKILLS.md +
skills/learned/β conventions that stuck - sessions/ β resumable snapshots
Tomorrow's run starts smarter than today's. That's the product. (Also: please donβt rm -rf .slmcode for sport.)
07 Permissions are a feature π‘οΈ¶
| Mode | Use when |
|---|---|
auto | You trust the loop (or it's a playground) π |
dry-run | Demos, CI dry checks, βwhat would you do?β π |
review | Real repos β stage patches, then slmcode apply π |
Shell is separate: shell_permission: allow | ask | deny. Files and shells have different blast radii. Treat them that way.
08 Any LLM, same loop π¶
Providers are adapters. The harness stays constant.
- π Local SLM β more
think_passes, smallermax_context_kb, patience - βοΈ Frontier β raise parallel, enjoy speed, keep inspectability
Next πΊοΈ¶
- β±οΈ Quick start β feel it
- π§ User guide β drive it daily
- ποΈ Architecture β package map for contributors
βοΈ Made with β₯ by UnicoLab