Imported from g-arruda/monetary_shocks_asset_prices (
arquivo/codex/AGENTS.md). Install upstream withnpx skills add g-arruda/monetary_shocks_asset_prices --skill codex. Copyright stays with the author.
Repository Guidelines
Project and authoritative context
This project adapts Alessi and Kerssenfischer's large-dimensional DFM identification strategy to Brazilian monetary shocks and asset prices. Before proposing methodology or interpreting results, read README.md, registro/metodo.md, registro/pendencias.md, and registro/historico_decisoes.md. The decision history is mandatory: do not revive rejected specifications or re-derive closed questions without new evidence. Use the dated working notes and generated reports as provenance, and heed banners marking an analysis as superseded. notas/_indice.md names the current DFM production note and flags which artifacts still belong to an older vintage. registro/pendencias.md was zeroed on 2026-09-08 by author decision (scope reset, not a verdict on any item); its cut items remain recoverable via git log -p -- registro/pendencias.md. email/ holds the running, verbatim email exchange with the advisor, one file per message.
The current paper is paper/paper_anpec.tex. arquivo/tex/ is an archived prose source, not an active manuscript. The written results source is output/irf/irf_section.md; confirm numerical claims against the underlying CSV/RDS artifacts. As of 2026-09-08, the paper and irf_section.md still describe the prior 115-series, 2013-01--2025-09 production window at (r,q,p)=(4,4,4) — one vintage behind the code in window and, outside the methodology, in inference (see below). On 2026-09-16 the inference half was closed in prose only: §3.5 dropped the wild bootstrap, §3.6 states the Anderson–Rubin sets and, in prose, that only the vectors of the identified ratio change, with a footnote putting MOSW's equation beside the DFM's, and sec:weak_iv no longer contrasts the two inferences, while §4 and the two wild bootstrap captions still describe the figures as painted; do not treat their numbers as current without checking notas/_indice.md first. script/fig_section5.R and script/fig_weak_iv.R are frozen behind --repaint-paper-figures for exactly that reason.
Pipeline
Run from the repository root. The main order is:
script/download.Rwritesdata/raw/raw_data.csv.script/clean.Rwritesdata/processed/data_log_deseasonalized.csv.script/instrument.Rbuilds the three monthly instrument variants oftab:first_stagethroughR/instrument/build_variants.R.script/model_alessi.Restimates the DFM;script/model_var.Ris the small-VAR benchmark.script/run_all.Rorchestrates declared stages and prerequisites.
Read script/README.md before selecting diagnostics or validation scripts. Run only the relevant stage for local changes; expensive production runs require deliberate scope. There is no conventional unit-test suite, linter, or package build. Validation scripts, in script/validation/ since 2026-09-17, and their committed fixtures in output/validation/ are the executable checks: rerun them after touching R/modeling/ or R/identification/. Scripts of closed rounds were archived to arquivo/script/ on the same date and on 2026-09-18 (arquivo/README.md). Preserve their fail-loud behavior and inspect git diff --check plus explicit git status --short paths before committing.
Identification and model invariants
- The production instrument is
z_jk_bs_purif(production_spec()$instrument). Since 2026-09-18 only the three variants oftab:first_stageexist in live code —z_bruto,z_bs_purif,z_jk_bs_purif— and their exact construction is defined in the builder; do not reconstruct the others from prose or archived code. data/raw/yields/yields_dia.csvanddata/raw/CDS 5y.xlsxare fixed external inputs with no repository producer. Treat them as read-only. Do not claim the yield curve is reproducible from this repository.- The production monthly sample is 2012-03 through 2025-12 (166 months, 162 factor innovations after
p=4) with 115 series (drop_setor_externo__eua__credito__imoveis_fiscal_expectations; paths, dimensions and samples read fromR/modeling/production_spec.R, which since 2026-09-18 carries only what two or more scripts read); the 106-series base is retained only for historical factor-grid reproduction. Production uses(r,q,p)=(5,5,4), withr=5selected by the BLL Bai–Ng surface (r=1,...,20: IC1=IC2=5, IC3=20),q=r=5as the operational decision, andp=4inherited from the prior vintage (an AIC check on the extended common sample,T=154, still picksp=4at 8.231267; BIC still picksp=2). The estimated factor VAR remains intercept-only. The Lenza-Primiceri COVID-volatility scale in the factor VAR (2026-09-14) is implemented but off:production_spec()$covid_volatilityis NULL, and with it NULL the whole production object isidentical()to the untreated code. It reweights the estimation, identifies nothing, and is not the abandoned heteroskedasticity identification; its treated sets condition on θ̂. The policy normalization variable isyield_6m, with a +50 bp impact shock. The frozen cell has ξ_mp = 6.057014 full / 8.643436 pre-COVID, F_rob,mp = 9.625428 / 13.809985, both companions stable (0.970090 / 0.993359). - The only active small-VAR benchmark is
ibc5_fx_cds_level_trend_p2. It usesibc_br,price_ipca,yield_6m,cambio_usd, andcds_5yin levels, with a constant and linear trend in every equation. AIC and BIC use a common 141-observation sample; production usesp=2, selected by AIC. Its published responses are horizon-specificC_h B_1estimates with VAR-only Anderson--Rubin/MOSW sets at 68% and 90% under NW(0). Do not add sensitivity cells, sum responses across horizons, or restore bootstrap inference for this benchmark. - Use
res$irfs, notres$irf, and recover variable names from the estimation data. Preserve the documented factor-space dimensions and normalization when comparing IRFs. - Factor selection uses the BLL-standardized Bai–Ng/Amengual–Watson variants. Plain Bai–Ng (2002) is inappropriate because the panel is non-stationary by design.
- Weak-IV conclusions are governed by the factor-space MOSW statistics, not by legacy first-stage rulers alone. Since 2026-09-08 the operational DFM inference is the 68%/90% Anderson–Rubin sets by test inversion, by author decision reversing the 2026-08-12 withdrawal (
registro/historico_decisoes.md§7).production_spec()$inferenceis the single authority andcompute_irf_dfm(inference=)the single switch ("ar", or"none"for the point alone); the wild bootstrap and the Kilian correction were removed from the code on 2026-09-18, andscript/ar_bands.R, which compared against them, was archived toarquivo/script/. Two properties travel with the sets and must be stated, not hidden: the plug-in covariance conditions on the estimatedΛ,K,Mandsy; andmosw_rform_covrequireshac_dim < T, which blocks(r,q)=(8,8)atp=4(336 ≥ 162) and the entire pre-COVID window atp=4(135 ≥ 90) — reported blocked, never patched with a pseudo-inverse or a substitute bootstrap.hac_dim < Tis this project's safeguard, not the authors' condition:CovAhat_Sigmahat_Gamma.m:91-95statesn²p + n(n+1)/2 + nk < Ton the parameter dimension (120 against 135 in the production shape), which is the necessary one sincerank(WHat) ≤ min(par_dim, T−1); the moment-dimension gate is sufficient and strictly stronger. Since the 2026-09-08 fidelity audit both are checked, and both bar exactly the same two cells, so no published number depends on which binds. Significance means "the set excludes zero", read off the topology throughar_excludes_zero():two_raysexcludes zero only when zero lies in its gap,real_linenever does,emptyis a misspecification signal. The set is bounded at level κ iff ξ_mp > κ. The small-VAR benchmark keeps its own AR/MOSW sets from the same module (Load/Innerdefault to the identity, andscript/validation/validate_mosw_ar.Rguards that the VAR path stays bit-identical). Those sets are inference for the VAR, never for the DFM, and weak-IV robustness is not instrument validity — both models use the same proxy. Do not change confidence-band interpretation or attribution without checking the current reports and pending-items file.
For changes to the estimation core, reproduce the current impact smoke test documented in the source guidance or current validation scripts before accepting results. Compare at least yield_6m, yield_2y, yield_5y, asset_ibov, and cambio_usd; do not update expected values merely to make a changed implementation pass.
Repository boundaries
R/contains reusable modules; nothing there should source a file fromscript/.script/contains entry points and diagnostics. Keep orchestration out of reusable modules.diagnostics/audits the production artifacts and must not mutate estimation code or production outputs.codigos_externos/is gitignored, read-only reference code. Production and validation paths must use project-owned implementations or committed fixtures.arquivo/is preserved historical material. No live path may source from it or write into it.data/is gitignored, in two levels:data/raw/is untreated download output and is never hand-edited;data/processed/is what goes into estimation.output/contains tracked estimation artifacts (exceptoutput/logs/); regenerate only those owned by the stage you ran and record the producing script.- The record splits three ways and the three never mix.
registro/is the living memory, edited in place:metodo.md(the design),pendencias.md(what is open),historico_decisoes.md(what died and why).notas/is the evidentiary record: one dated, append-only note per round, carrying a vintage banner — this is what the paper pulls numbers from.pareceres/is what outside reviewers sent in, kept verbatim.email/is the running email exchange with the advisor, one file per message (email_{meu,professor}_DD-MM_HHhMM.md), kept verbatim likepareceres/but ongoing rather than a single delivered document.progress_logs/is session continuity and is disposable. Never write a session log intonotas/, and never leave a durable result inprogress_logs/.
Coding, figures, and prose
Use English for code, identifiers, and code comments; Portuguese is appropriate in registro/, notas/, pareceres/, email/, reports, and the paper. Keep comments limited to non-obvious methodological choices. Use ggplot2 for paper figures and preserve the established shaded 80% and 90% band style. Fail clearly on missing inputs, dimension mismatches, and invalid numerical states. Never fabricate fallback data or silently substitute a different estimator.
When results change, distinguish a specification change from an implementation bug, update the relevant authoritative note/report, and keep generated prose synchronized with source tables. Do not describe 68% bands as statistical significance where the project's two-tier reading rule reserves “significativo” for the 90% band.
Git hygiene
The worktree often contains active research changes and regenerated artifacts. Preserve unrelated modifications, stage explicit paths only, and never use broad staging. Commit messages should describe the research, identification, code, or result change. Do not add AI attribution.