Imported from BauplanLabs/bauplan-skills (
AGENTS.md). Install upstream withnpx skills add BauplanLabs/bauplan-skills. Copyright stays with the author.
Bauplan
Git Workflow (code repo)
This section is about git branches for source code, not Bauplan data branches.
- Always create a new git branch before making any changes. Use a descriptive branch name based on the task (e.g.,
fix/login-bug,feat/add-search). - All work happens on feature branches. The
mainbranch is off-limits for direct edits. - Commit every change immediately after making it. Each commit should have a clear, one-line message describing what changed.
- Keep commits small and focused. One logical change per commit.
- When the task is complete, leave the branch ready for review. Do not merge into
main.
Bauplan Workflow (data lakehouse)
Bauplan is a data lakehouse platform where data changes follow a Git-like workflow, but on Bauplan data branches (managed by bauplan branch), not git branches. A typical workflow looks like this:
- Create and switch to a new data branch:
bauplan checkout -b <branch> --from-ref main - Iterate on your pipeline code and test with
bauplan run --dry-run - When ready, run
bauplan runto materialize — the code is snapshotted by Bauplan and saved under the jobId the platform returns - Review changes with
bauplan branch diff main - Merge into
mainto publish
Hard safety rules (always)
- Never publish by writing directly on
main. Use a Bauplan branch and merge to publish. - Never import data directly into
main. - Before merging into
main, runbauplan branch diff mainand review changes, or usebauplan queryover the two branches to quickly compare how data changes in the target tables. - Prefer
bauplan run --dry-runduring iteration because it is much faster and safer. However, tables will not be materialized in the lake during a dry run, so the only way to preview table content is to use--preview(e.g.,bauplan run --dry-run --preview head). Materialization is blocked onmain. Every branch is readable, but only branches prefixed with your username (username.branchname) are writable by you. - When handling external API keys (LLM keys), do not hardcode them in code or commit them. Use Bauplan parameters or secrets.
If any instruction conflicts with these rules, the rules win.
CLI vs Python SDK
Use the CLI for interactive exploration, quick inspections, and one-off commands:
bauplan table get <namespace>.<table>— inspect table metadatabauplan query "<sql>"— run a querybauplan branch ls,bauplan run --dry-run, etc.
Important: By default, bauplan query returns only 10 rows (--max-rows 10). Use --all-rows to retrieve the full result set. For large outputs, pipe to a local file or use --output json to get machine-readable output (e.g., bauplan query --all-rows --output json "SELECT ..." > results.json).
Use the Python SDK when:
- You need to process or transform large result sets —
client.query()returns a fullpyarrow.Tablewith no row limit by default - A Python script is more natural than a sequence of shell commands
- You need programmatic control (loops, conditionals, error handling)
- You are writing pipelines, ingestion scripts, or automation
General Python guidance
- Use
uvto run Python scripts and manage dependencies (e.g.,uv run python3 script.py). - Verify that generated Python compiles and passes lint with
uvx ruff check,uvx ruff format, anduvx ty check. - Do not guess flags or method names. If you get stuck or need method signatures, use
WebFetchto pull the relevant markdown page fromhttps://docs.bauplanlabs.com/llms.txt(see "Looking up documentation" below).
Bauplan Python client
import bauplan
# Default: authenticates from BAUPLAN_API_KEY >> BAUPLAN_PROFILE >> ~/.bauplan/config.yml
client = bauplan.Client()
# Or specify a profile explicitly
client = bauplan.Client(profile='default')
The client returns Arrow tables. Do not use pandas: Polars has zero-copy Arrow interop and is faster. Common patterns:
import polars as pl
# Query → Arrow table
table = client.query('SELECT * FROM ns.my_table', ref='my_branch')
# Arrow table → Polars DataFrame (zero-copy)
df = pl.DataFrame(table)
To get the current username (e.g., for branch naming):
bauplan info
Looking up documentation
Bauplan publishes an LLM-friendly documentation index at https://docs.bauplanlabs.com/llms.txt. This file lists every doc page as a markdown URL (e.g., https://docs.bauplanlabs.com/concepts/models.md). Use WebFetch to pull any page directly: the markdown format is much more reliable than web searching.
Key pages by topic:
| Topic | URL |
|---|---|
| Python SDK reference | https://docs.bauplanlabs.com/reference/bauplan.md |
| SDK types | https://docs.bauplanlabs.com/reference/bauplan-sdk-types.md |
| Catalog types | https://docs.bauplanlabs.com/reference/bauplan-schema.md |
| CLI reference | https://docs.bauplanlabs.com/reference/cli.md |
| Standard expectations | https://docs.bauplanlabs.com/reference/bauplan-standard-expectations.md |
| Models | https://docs.bauplanlabs.com/concepts/models.md |
| Semantic annotations | https://docs.bauplanlabs.com/concepts/semantic-annotations.md |
| Pipelines | https://docs.bauplanlabs.com/concepts/pipelines.md |
| Tables | https://docs.bauplanlabs.com/concepts/tables.md |
| Namespaces | https://docs.bauplanlabs.com/concepts/namespaces.md |
| Expectations | https://docs.bauplanlabs.com/concepts/expectations.md |
| Data branches | https://docs.bauplanlabs.com/concepts/git-for-data/data-branches.md |
| Import data | https://docs.bauplanlabs.com/tutorial/import.md |
| Schema conflicts | https://docs.bauplanlabs.com/common-scenarios/schema-conflicts.md |
| Secrets | https://docs.bauplanlabs.com/common-scenarios/parameterized-runs.md |
| Parameters | https://docs.bauplanlabs.com/common-scenarios/parameterized-runs.md |
| Execution model | https://docs.bauplanlabs.com/overview/execution-model.md |
When unsure about a method, flag, or concept, fetch the relevant page rather than guessing. For the full index: https://docs.bauplanlabs.com/llms.txt
CLI: The bauplan CLI is also self-documenting:
bauplan --help— lists all available commandsbauplan <command> --help— shows arguments and options for a specific command (e.g.,bauplan query --help,bauplan branch --help)
Embedded data documentation
In Bauplan, table-level and column-level documentation is accessible by the CLI and Python SDK. This documentation can hold any text the user would like, but it is best to use as a semantic layer that describes the intent and context of tables and columns to ensure correct usage of the data. It is stored in Iceberg and versioned per data branch, so a table documented on one branch may be bare on another.
With the CLI. When the bauplan table get ... command is used, table-level documentation is first shown with the title Table Documentation. Column-level documentation is available in the DOC column, although it is truncated if longer than a fixed amount or if it spans more than one line. If the column-level documentation is truncated, a note is printed that will direct the user to use -O json to receive the full documentation. In the JSON output, table-level documentation is at properties.comment rather than a top-level key, and column-level documentation is at fields[].doc.
With the Python SDK. When bauplan.Client.get_table is called, a bauplan.schema.Table instance is returned; the table-level documentation is accessible by the bauplan.schema.Table.comment property. The column-level documentation is accessible by the bauplan.schema.TableField.doc property, which is never truncated.
Authoring. Documentation is only written for materialized models when running a pipeline. Table-level documentation is set from the model's function docstring if defined, or from the output schema's class docstring otherwise. Column-level documentation is set from the doc parameter of the TableField annotation. Both docstrings pass through inspect.cleandoc, so indentation is normalized and a whitespace-only docstring means no documentation at all.
Skills
Skills whose names start with bauplan- contain use-case-specific instructions (e.g., building pipelines, ingesting data, debugging failed runs); when a relevant skill is available, follow its guidance for that workflow.
Authentication
Assume Bauplan credentials are available via local CLI config, environment variables, or a profile. Do not ever prompt for API keys, nor ask the user to tell you their API key: if there are no keys set, tell the user to visit https://app.bauplanlabs.com/dashboard, get the key and do the setup following the instructions on the screen.