cristoslc/architecture-reference · Archived

discover-architecture

Analyze a codebase to discover its architecture style(s) through direct source code inspection. Reads actual code structure, dependency graphs, module boundaries, and runtime patterns — no heuristics or filesystem signal counting. Produces a classification report with evidence citations. Use when the user says "what architecture is this", "analyze this repo", "discover architecture", "classify this codebase", "what patterns does this use", or points you at a repo and asks about its structure. A…

First seen Mar 9, 2026

Installation

$ npx skills add cristoslc/architecture-reference --skill discover-architecture

Summary

  • Analyze a codebase to discover its architecture style(s) through direct source code inspection.
  • Reads actual code structure, dependency graphs, module boundaries, and runtime patterns — no heuristics or filesystem signal counting.
  • Produces a classification report with evidence citations.
  • Use when the user says "what architecture is this", "analyze this repo", "discover architecture", "classify this codebase", "what patterns does this use", or points you at a repo and asks about its structure.
  • Also triggers on "architecture review", "codebase analysis", or "how is this system structured".

Stronger alternatives

This repository is archived — consider an actively maintained alternative.

Similar popular skills

Related neighbors and high-traction skills in the same topics — useful to compare before installing.

Also in this package

Other skills from cristoslc/architecture-reference.

npx skills add cristoslc/architecture-reference

Browse all from cristoslc/architecture-reference

More details

Agent compatibility

Declared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.

Claude Code Not declared
Cursor Not declared
Codex Not declared
GitHub Copilot Not declared
Windsurf Not declared
Gemini CLI Not declared
Cline Not declared
OpenCode Not declared

Repository health

Stars 1
License MIT
Default branch main
Open issues 0
Status Archived

Skill metadata

Parsed from SKILL.md frontmatter.

Version2.0.0
LicenseMIT
Allowed toolsBash, Read, Grep, Glob, Agent
More metadata
short-description
Discover architecture styles through deep source code analysis
version
2.0.0
author
cristos
source-repo
https://github.com/cristoslc/architecture-reference

Package contents

Files included with this skill beyond the listing page.

  • skill md SKILL.md 12,898 B
  • docs SUMMARY.md 625 B

History

  1. First seen on skills.sh
  2. First recorded snapshot · 1 installs

SKILL.md

Discover Architecture

Analyze a codebase to identify its architecture style(s) by reading actual source code — not by counting filesystem signals or matching directory name patterns. Architecture is about structure and relationships: how modules communicate, where boundaries are enforced, what the dependency graph looks like. These things require reading code, not scanning for Dockerfiles.

The 12 Canonical Styles

These are the only valid architecture style classifications. Read references/styles.md for full definitions, distinguishing characteristics, and production frequency data.

Style What to look for
Microkernel Host application with plugin/extension registry. Core provides lifecycle management; plugins provide domain behavior. Extension point contracts, dynamic module loading.
Layered Horizontal separation into layers (presentation/business/data). Strict dependency direction — upper layers depend on lower, never reverse.
Modular Monolith Single deployable unit with well-defined module boundaries. Modules are logically independent but physically coupled. Module registries, feature toggles.
Event-Driven Components communicate through events, not direct calls. Message brokers, event buses, pub/sub patterns, async handlers.
Pipeline Data flows through ordered processing stages. Each stage transforms input to output. Middleware chains, filter pipelines, compiler passes.
Microservices Independent services with own databases, deployed separately. Service mesh, API gateways, per-service CI/CD.
Service-Based Coarse-grained services sharing infrastructure. Less distributed than microservices — shared databases, simpler communication.
Hexagonal Architecture Core business logic isolated in center. External concerns connect through ports (interfaces) and adapters (implementations). Dependency inversion enforcement.
Domain-Driven Design Code organized around business domains (bounded contexts). Aggregates, domain events, ubiquitous language, repository pattern.
Multi-Agent Multiple autonomous agents with specialized capabilities collaborating through message passing or supervisor hierarchies.
Space-Based In-memory distributed data grid with peer-to-peer replication. Masterless, eventual consistency.
CQRS Separate read and write models. Command/query separation, event stores, projection builders.

How to Classify

Step 1: Inventory the codebase

Get oriented quickly. Run these in parallel:

  • ls the root directory and key subdirectories (src/, lib/, packages/, services/, internal/, cmd/)
  • Read README.md (first 200 lines) — often states what the project IS
  • Read any ARCHITECTURE.md, docs/architecture/, or CONTRIBUTING.md
  • Check package metadata (package.json description, pyproject.toml, pom.xml, Cargo.toml, go.mod)
  • Note the primary language(s) and framework(s)
  • Check for deployment configs: docker-compose.yaml/compose.yaml, Dockerfile(s), k8s//helm/, nginx.conf/reverse proxy configs, serverless.yml, fly.toml, Procfile, terraform/ — these define runtime components, boundaries, and contracts that are architectural evidence on equal footing with application code

Step 2: Inspect code structure

This is where classification happens. Read actual source files — not just directory names.

For each candidate style, look for structural evidence:

  • Module boundaries: How is code organized? Are there clear module interfaces, or is everything in one flat namespace?
  • Dependency direction: Do dependencies flow in one direction? Is there dependency inversion?
  • Communication patterns: How do components talk to each other? Direct function calls? Message passing? HTTP? Events?
  • Extension mechanisms: Are there plugin registries, middleware chains, hook systems?
  • Data flow: Does data flow through transformation stages, or is it request/response through layers?
  • Deployment topology: Is this one deployable unit or many? How do you tell? Do deployment configs introduce runtime components, boundaries, or contracts not visible in application code?

Read at least:

  • 2-3 "entrypoint" files (main.go, app.py, Program.cs, index.ts, etc.)
  • The dependency injection / wiring configuration (if any)
  • 2-3 representative domain files showing the core architecture
  • Any inter-module or inter-service communication code
  • Deployment configs found in Step 1 (compose files, proxy configs, k8s manifests) — read these as architecture, not ops. A compose file can define components (a reverse proxy container), boundaries (which container is network-facing), and contracts (a shared volume between producer and consumer). When a deployment config introduces any of these, it is architectural evidence for the style classification.

When is deployment topology architecturally significant? Not every Dockerfile matters. The test: does the deployment config introduce a component, boundary, or contract that is invisible in application source code? A single Dockerfile that packages the app is ops. A compose stack where nginx is the sole network-facing surface and a named volume is the inter-component contract — that's architecture. Report deployment-level findings inline with the style they inform (e.g., a two-container split within a Modular Monolith), not as a separate dimension.

Step 3: Classify with evidence

Based on what you read, determine:

  1. Primary style(s) — the dominant architectural pattern(s). Most production repos exhibit 2 styles (74% in the evidence base).
  2. Confidence — how clear the evidence is (0.0-1.0)
  3. Evidence citations — specific files, classes, and patterns that support each classification

Classification principles:

  • Classify what the repo IS, not what it enables. A plugin framework IS Microkernel. A message broker IS Event-Driven infrastructure.
  • Multi-style composition is normal. Don't force single-style classification. A system can be both Layered and Microkernel (layered internal structure with plugin extension).
  • Code trumps documentation. If the README says "microservices" but the code is a monolith with a single database, classify based on code.
  • Distinguish style from technology. Having Docker doesn't make it Microservices. Having Kafka doesn't make it Event-Driven. Look at how the architecture is actually structured.
  • Deployment configs are source files. A compose file that introduces runtime components (a reverse proxy, a worker, a sidecar) or defines inter-component contracts (shared volumes, network boundaries) is architectural evidence. Fold deployment-level findings into the style they inform — don't create a separate "deployment architecture" section. If the deployment topology contradicts the code-level architecture (code looks like a monolith but deploys as independent services, or vice versa), note the tension.
  • If truly indeterminate, say so and explain why. Don't guess.

Step 4: Determine scope and use-type

Scope (per ADR-001):

  • Platform: designed to be extended/built upon — plugin systems, API surfaces, infrastructure for other software (e.g., Kafka, Grafana, VS Code)
  • Application: end-user facing, solves a specific problem (e.g., Mastodon, Ghostfolio)

Use-type (per ADR-001):

  • Production: real system used in production or production-ready
  • Reference: educational, demo, template, starter kit

Step 5: Identify quality attributes

Only report quality attributes with direct evidence in code:

QA Evidence to look for
Deployability Container configs, CI/CD pipelines, deployment scripts
Modularity Clear module boundaries, DI configuration, interface segregation
Scalability Horizontal scaling configs, sharding, message queues
Fault Tolerance Circuit breakers, retry policies, health checks, graceful degradation
Observability Structured logging, metrics, tracing (OpenTelemetry, Prometheus)
Evolvability Plugin systems, extension points, feature flags

Do NOT report quality attributes you can't see in code (Performance, Security, Testability, etc.) — these are real but invisible in source analysis. Note this limitation in the report.

Step 6: Infer domain

From README, package metadata, directory naming, and code content. Use specific domains when clear (E-Commerce, Developer Tools, Observability, etc.). Use "General Purpose" if unclear.

Output Format

Every classification produces two outputs from the same analysis pass:

  1. Report (references/report.template.j2): human-readable markdown with style rationales, evidence citations, quality attributes. Saved to docs/architecture-reports/<project-name>-<YYYY-MM-DD>.md.
  2. Catalog entry (references/catalog-entry.template.j2): machine-consumable YAML for pipeline ingestion and the evidence base. Saved alongside the report or returned inline when used in batch mode.

Both templates share the same variable namespace — populate one set of variables and render both. The report is the authoritative analysis; the catalog entry is a structured projection of it.

Report guidelines

The report should fit on roughly one printed page — if you're writing paragraphs, you're writing too much. Specific length targets:

  • Style rationale: 2-3 sentences. Lead with the structural pattern, cite 1-2 key files, stop. No function-level walkthroughs.
  • Evidence table: Comma-separated file paths or short phrases in the Key Evidence column. Not prose.
  • Quality attributes: One line per QA. Format: Name: file1, file2, brief note. No full sentences.
  • Domain: One line.
  • Production context: 2-4 bullet points, each one fact with a number.

Example (correct length):

### Pipeline (primary)

Data flows through five ordered stages in `generate.sh`: fetch → clean → score → compress → render. Each stage writes to `fetched/<timestamp>/` and the next stage reads from it. Partial re-runs (`./generate.sh linear`) confirm stage composability.

Not this (too long — don't trace every function call):

### Pipeline (primary)

The entire system is structured as an ordered chain of data-transformation
stages with clear input/output contracts at each boundary. `generate.sh`
drives five sequential stages: (1) Fetch — independent per-source scripts
(`sources/linear/fetch.py`, `sources/quickbase/fetch.py` ...) [300 more words]

Catalog entry guidelines

The catalog entry captures the classification result in a flat YAML structure consumable by pipeline tooling. Key fields:

  • architecture_styles: flat list of style names (primary first, then secondary)
  • classification_confidence: the overall confidence score (0.0–1.0)
  • classification_reasoning: the full analysis text — what was found, what was ruled out, and why
  • onelinesummary: a single sentence capturing the project's architecture in plain language

Pipeline scripts inject mechanical metadata at write time (classificationmodel, classificationmethod, classificationdate, discoveredat, etc.) — the classifying agent does not need to fabricate these.

Edge Cases

Monorepo with multiple projects: Note the monorepo structure. Classify the overall architecture, noting sub-projects if they have distinct styles.

Trivial repos (< 10 source files): Classification may not be meaningful. Say so.

Libraries/frameworks (consumed as dependencies, not deployed): These have no deployable architecture. Classify as what they ARE (a plugin framework is Microkernel), note they're libraries.

Indeterminate: If code is too flat, too small, or too unconventional to classify, say "Indeterminate" and explain what would help (more code, clearer boundaries, etc.).