Repository Analysis

dathere/qsv

Blazing-fast Data-Wrangling toolkit

3.7 Likely human-written View on GitHub

Analysis Overview

This report presents the forensic synthetic code analysis of dathere/qsv, a Rust project with 3,719 GitHub stars. SynthScan v2.0 examined 395,989 lines of code across 753 source files, recording 909 pattern matches distributed across 18 syntactic categories. The overall adjusted score of 3.7 places this repository in the Likely human-written band.

The scanner applied 160+ deterministic lexical heuristics, multi-line block detectors, abstract syntax tree depth profilers, and a cross-file Jaccard similarity matrix to construct a statistically normalised synthetic code estimate. All matches are individually weighted by severity coefficient and contextual multiplier before summation, and the resulting headline score is temporally discounted to account for the repository's development history relative to the commercial emergence of large language model coding tooling (November 2022 onward).

3.7
Adjusted Score
3.7
Raw Score
100%
Time Factor
2026-07-14
Last Push
3.7K
Stars
Rust
Language
396.0K
Lines of Code
753
Files
909
Pattern Hits
2026-07-14
Scan Date
0.02
HC Hit Rate

What These Metrics Mean

Adjusted Score
Primary synthetic code indicator. Raw score normalised per 1,000 lines of code and multiplied by the temporal discount factor. This is the definitive comparative metric — use it to rank repositories by AI authorship density.
Raw Score
The unmodified sum of all severity-weighted, context-multiplied pattern match scores before temporal discounting. Reflects the absolute signal strength independent of when the repository was last active.
Time Factor
The temporal discount multiplier (0–100%) applied to the raw score. Repositories last updated before ChatGPT's launch (Nov 2022) receive a 5% factor. Full signal is only assigned to repositories active in the post-adoption era (Jan 2024+).
Pattern Hits
Total count of individual pattern matches across all files and categories. A high hit count with a low score may indicate a very large codebase with isolated AI snippets; a low count with a high score indicates dense, concentrated AI signatures.
HC Hit Rate
High+Critical pattern hits per file, averaged across the repository. This orthogonal signal catches repositories where a few files are densely packed with high-severity AI tells — a strong indicator even when the normalised score appears moderate due to codebase size.
Lines of Code / Files
Total lines and files analysed. The scanner examines 94 file extensions. These denominators are used to normalise the score, enabling fair comparison between repositories of vastly different sizes.

Score History

This chart maps the temporal evolution of the adjusted synthetic code score across successive scan runs. An upward trajectory indicates ongoing incorporation of AI-generated code or expanding LLM-assisted scaffolding; a stable or declining trajectory may reflect active human refactoring, code removal, or the adoption of stricter authorship policies. The dashed secondary line (right axis) independently tracks total raw pattern hit count, which can diverge from the normalised score when codebase size changes significantly between scans.

Severity Breakdown

Classifies detected patterns by their diagnostic confidence and structural impact. CRITICAL patterns (coefficient 10) represent definitive synthetic signatures — hallucinated imports, explicit LLM attribution metadata — virtually never produced by human authors. HIGH (5) indicates strong structural tells such as cross-file repetition or cross-linguistic idioms. MEDIUM (2) covers recognisable conversational padding and AI-specific vocabulary. LOW (1) captures subtle indicators like tautological comments and generic boilerplate that require density to carry independent signal.

CRITICAL 0HIGH 14MEDIUM 227LOW 668

Directory Score Breakdown

This horizontal bar chart decomposes the repository's raw synthetic code score by top-level directory, allowing you to pinpoint precisely which modules or components carry the highest AI authorship density. Directories with disproportionately high scores relative to their size warrant targeted manual review: concentrated AI signatures often trace back to mass-generated configuration layers, auto-ported test suites, LLM-scaffolded boilerplate classes, or entire subsystems authored under heavy copilot assistance. Use this view to prioritise your human code-review effort.

Pattern Findings

The scanner identified 909 distinct pattern matches across 18 syntactic categories. Each entry below represents a discrete location in the source code where the engine recorded a statistically significant AI authorship indicator. Expand any category row to inspect the individual file paths, line numbers, code snippets, and the lexical context (CODE, COMMENT, or STRING) in which each match was detected.

Reading the findings table: The Severity column indicates the diagnostic confidence level (CRITICAL / HIGH / MEDIUM / LOW). The Context column identifies whether the match occurred inside executable code, an inline comment, or a string literal — comment-context matches receive a ×1.5 weight because LLMs systematically over-annotate. The ⚡ bolt icon marks clustered matches: three or more patterns within a 10-line window, each receiving an additional ×1.5 density multiplier as dense clusters constitute far stronger evidence of synthetic authorship than isolated hits.

Decorative Section Separators163 hits · 510 pts
SeverityFileLineSnippetContext
MEDIUMCargo.toml527# ================================COMMENT
MEDIUMresources/describegpt_md_defaults.toml230# ─────────────────────────────────────────────────────────────────────────────STRING
MEDIUMresources/profiles/geoconnex.yaml30# =========================================================================COMMENT
MEDIUMresources/profiles/geoconnex.yaml32# =========================================================================COMMENT
MEDIUMresources/profiles/geoconnex.yaml59# =========================================================================COMMENT
MEDIUMresources/profiles/geoconnex.yaml61# =========================================================================COMMENT
MEDIUMresources/profiles/geoconnex.yaml83# =========================================================================COMMENT
MEDIUMresources/profiles/geoconnex.yaml85# =========================================================================COMMENT
MEDIUMresources/profiles/geoconnex.yaml121# =========================================================================COMMENT
MEDIUMresources/profiles/geoconnex.yaml123# =========================================================================COMMENT
MEDIUMresources/profiles/geoconnex.yaml229# =========================================================================COMMENT
MEDIUMresources/profiles/geoconnex.yaml231# =========================================================================COMMENT
MEDIUMresources/profiles/geoconnex.yaml260# =========================================================================COMMENT
MEDIUMresources/profiles/geoconnex.yaml262# =========================================================================COMMENT
MEDIUMresources/profiles/geoconnex.yaml282# =========================================================================COMMENT
MEDIUMresources/profiles/geoconnex.yaml284# =========================================================================COMMENT
MEDIUMresources/profiles/dcat-ap-v3.yaml21# =========================================================================COMMENT
MEDIUMresources/profiles/dcat-ap-v3.yaml23# =========================================================================COMMENT
MEDIUMresources/profiles/dcat-ap-v3.yaml92# =========================================================================COMMENT
MEDIUMresources/profiles/dcat-ap-v3.yaml94# =========================================================================COMMENT
MEDIUMresources/profiles/dcat-ap-v3.yaml131# =========================================================================COMMENT
MEDIUMresources/profiles/dcat-ap-v3.yaml133# =========================================================================COMMENT
MEDIUMresources/profiles/dcat-ap-v3.yaml182# =========================================================================COMMENT
MEDIUMresources/profiles/dcat-ap-v3.yaml184# =========================================================================COMMENT
MEDIUMresources/profiles/dcat-ap-v3.yaml305# =========================================================================COMMENT
MEDIUMresources/profiles/dcat-ap-v3.yaml307# =========================================================================COMMENT
MEDIUMresources/profiles/dcat-ap-v3.yaml395# =========================================================================COMMENT
MEDIUMresources/profiles/dcat-ap-v3.yaml397# =========================================================================COMMENT
MEDIUMresources/profiles/dcat-ap-v3.yaml405# =========================================================================COMMENT
MEDIUMresources/profiles/dcat-ap-v3.yaml407# =========================================================================COMMENT
MEDIUMresources/profiles/croissant.yaml39# =========================================================================COMMENT
MEDIUMresources/profiles/croissant.yaml41# =========================================================================COMMENT
MEDIUMresources/profiles/croissant.yaml68# =========================================================================COMMENT
MEDIUMresources/profiles/croissant.yaml70# =========================================================================COMMENT
MEDIUMresources/profiles/croissant.yaml93# =========================================================================COMMENT
MEDIUMresources/profiles/croissant.yaml95# =========================================================================COMMENT
MEDIUMresources/profiles/croissant.yaml119# =========================================================================COMMENT
MEDIUMresources/profiles/croissant.yaml121# =========================================================================COMMENT
MEDIUMresources/profiles/croissant.yaml249# =========================================================================COMMENT
MEDIUMresources/profiles/croissant.yaml251# =========================================================================COMMENT
MEDIUMresources/profiles/croissant.yaml313# =========================================================================COMMENT
MEDIUMresources/profiles/croissant.yaml315# =========================================================================COMMENT
MEDIUMresources/profiles/croissant.yaml510# =========================================================================COMMENT
MEDIUMresources/profiles/croissant.yaml512# =========================================================================COMMENT
MEDIUMresources/profiles/croissant.yaml524# =========================================================================COMMENT
MEDIUMresources/profiles/croissant.yaml526# =========================================================================COMMENT
MEDIUMresources/profiles/dcat-us-v3.yaml22# =========================================================================COMMENT
MEDIUMresources/profiles/dcat-us-v3.yaml24# =========================================================================COMMENT
MEDIUMresources/profiles/dcat-us-v3.yaml112# =========================================================================COMMENT
MEDIUMresources/profiles/dcat-us-v3.yaml114# =========================================================================COMMENT
MEDIUMresources/profiles/dcat-us-v3.yaml182# =========================================================================COMMENT
MEDIUMresources/profiles/dcat-us-v3.yaml184# =========================================================================COMMENT
MEDIUMresources/profiles/dcat-us-v3.yaml208# =========================================================================COMMENT
MEDIUMresources/profiles/dcat-us-v3.yaml210# =========================================================================COMMENT
MEDIUMresources/profiles/dcat-us-v3.yaml422# =========================================================================COMMENT
MEDIUMresources/profiles/dcat-us-v3.yaml424# =========================================================================COMMENT
MEDIUMresources/profiles/dcat-us-v3.yaml531# =========================================================================COMMENT
MEDIUMresources/profiles/dcat-us-v3.yaml533# =========================================================================COMMENT
MEDIUMresources/profiles/dcat-us-v3.yaml541# =========================================================================COMMENT
MEDIUMresources/profiles/dcat-us-v3.yaml543# =========================================================================COMMENT
103 more matches not shown…
Over-Commented Block492 hits · 455 pts
SeverityFileLineSnippetContext
LOWCargo.toml21 "LICENSE-MIT",COMMENT
LOWCargo.toml461# https://github.com/plotly/plotly.rs/pull/425), a strict superset of the earlier pins:COMMENT
LOWCargo.toml481# pulled in by a NON-optional dependency edge, not by a feature we control —COMMENT
LOWCargo.toml501COMMENT
LOWCargo.toml521# (e.g. py-1.19.0)COMMENT
LOWCargo.toml621# evaluation; polars (SQL context) backs the SQL-requiring helpersCOMMENT
LOWbacon.toml21 "cargo", "clippy",COMMENT
LOW.serena/project.yml1# the name by which the project can be referenced within SerenaCOMMENT
LOW.serena/project.yml21# Note:COMMENT
LOW.serena/project.yml61- search_for_patternCOMMENT
LOW.serena/project.yml81base_modes:COMMENT
LOW.serena/project.yml101symbol_info_budget:COMMENT
LOW.serena/project.yml121# and cannot be accessed via read_memory or write_memory.COMMENT
LOW.serena/project.yml141# symbols and references across package boundaries.COMMENT
LOWresources/describegpt_defaults.toml21# otherwise insertsCOMMENT
LOWresources/describegpt_defaults.toml41# (the curated qsv vocabulary)COMMENT
LOWresources/describegpt_md_defaults.toml1name = "qsv Default Markdown Template File"COMMENT
LOWresources/describegpt_md_defaults.toml21# content_type, min, max, cardinality, enumeration, null_count,COMMENT
LOWresources/describegpt_md_defaults.toml101# the `{DATASET_DESCRIPTION}` and `{SEMANTICMD_TAGS}` placeholders are substituted afterwardsCOMMENT
LOWresources/profiles/geoconnex.yaml1# Geoconnex hydrologic linked-data projection profileCOMMENT
LOWresources/profiles/geoconnex.yaml21# Discovery merge: disabled. Geoconnex contributors typically authorCOMMENT
LOWresources/profiles/geoconnex.yaml121# =========================================================================COMMENT
LOWresources/profiles/geoconnex.yaml261# Catalog envelopeCOMMENT
LOWresources/profiles/dcat-ap-v3.yaml1# DCAT-AP v3 projection profile (https://semiceu.github.io/DCAT-AP/releases/3.0.0/)COMMENT
LOWresources/profiles/croissant.yaml1# Croissant 1.1 ML metadata profile (https://github.com/mlcommons/croissant)COMMENT
LOWresources/profiles/croissant.yaml21# remains the faster option but is not in the Croissant vocabulary;COMMENT
LOWresources/profiles/croissant.yaml241 # 1.1: `citeAs` is a first-class cr: term defined in the @context;COMMENT
LOWresources/profiles/croissant.yaml281 # `res.url` if one was supplied (URL inputs, --initial-context), elseCOMMENT
LOWresources/profiles/croissant.yaml301COMMENT
LOWresources/profiles/croissant.yaml321# build_croissant_fields helper) so we can include the canonicalCOMMENT
LOWresources/profiles/dcat-us-v3.yaml1# DCAT-US v3 projection profileCOMMENT
LOWtests/test_geocode.rs2261 // the input column is preserved & three dyncols columns are appendedCOMMENT
LOWtests/test_apply.rs1321 svec!["Just enter SSN when prompted. Also try SSN if it doesn't work."],COMMENT
LOWtests/test_profile.rs301COMMENT
LOWtests/test_profile.rs1641 // and lands at the default "stdin.metadata.json" in the workdir.COMMENT
LOWtests/test_profile.rs1921 );COMMENT
LOWtests/test_profile.rs2041#[test]COMMENT
LOWtests/test_sniff.rs741// Build a wide CSV (rows ~1KB each) of `n` data rows with columns id,code,pad.COMMENT
LOWtests/test_alloc_tuning.rs1//! Regression tests for startup allocator-tuning ordering (roborev #2717 / #2718).COMMENT
LOWtests/test_alloc_tuning.rs21COMMENT
LOWtests/test_synthesize.rs421 err.contains("does not exist") || err.contains("not a file"),COMMENT
LOWtests/test_synthesize.rs781 );COMMENT
LOWtests/test_applydp.rs401#[test]COMMENT
LOWtests/test_describegpt.rs3401 assert_eq!(COMMENT
LOWtests/test_stats.rs4521// The following tests verify the layered OOM-fallback behavior: whenCOMMENT
LOWtests/test_sample.rs2041// These spin up a local actix-web server with hand-built fixture bytes andCOMMENT
LOWtests/workdir.rs421 }COMMENT
LOWtests/test_split.rs1661// B2 — empty input + `--filter` no longer underflows in sequential split.COMMENT
LOWtests/test_sqlp.rs2341// let wrk = Workdir::new("sqlp_select_1");COMMENT
LOWtests/test_sqlp.rs3541// vec![COMMENT
LOWtests/test_sqlp.rs3561// )COMMENT
LOWtests/test_get.rs221COMMENT
LOWtests/test_get.rs261 // Path-style S3 object: object_store issues `GET /{bucket}/{key}`COMMENT
LOWtests/test_get.rs1341 .and_then(|n| n.to_str())COMMENT
LOWtests/test_to.rs261// let wrk = Workdir::new("to_sqlite_dir");COMMENT
LOWtests/test_to.rs281COMMENT
LOWtests/test_to.rs301COMMENT
LOWtests/test_to.rs321COMMENT
LOWtests/test_to.rs341// let places_iter = stmtCOMMENT
LOWtests/test_to.rs361// let wrk = Workdir::new("to_sqlite");COMMENT
432 more matches not shown…
AI Slop Vocabulary47 hits · 141 pts
SeverityFileLineSnippetContext
MEDIUMtests/test_select.rs463 // This should not panic with our new robust implementationCOMMENT
MEDIUMtests/test_select.rs487 // This should not panic with our new robust implementationCOMMENT
MEDIUMtests/test_synthesize.rs507 // "Tom"), so the "no real value leaked" assertion below is robust toCOMMENT
MEDIUMtests/test_stats.rs1191 // Parse the JSON and assert on fields (rather than raw substrings) so the test is robust toCOMMENT
MEDIUMtests/test_pivotp.rs1112 // Should use Median — outliers inflate stddev but MAD stays robustCOMMENT
MEDIUM.claude/skills/src/converted-file-manager.ts400 // Use path.relative() for robust cross-platform validation (handles Windows case-insensitivity)COMMENT
MEDIUM.claude/skills/src/converted-file-manager.ts725 // Use more robust parsing to handle edge cases like filenames containing ".converted."COMMENT
MEDIUM.claude/skills/src/converted-file-manager.ts1260 // Delete files with robust error handlingCOMMENT
MEDIUMcontrib/completions/src/usage_parser.rs5// completion generation. It uses the qsv_docopt parser for robust option typeCOMMENT
MEDIUMexamples/viz/smart_dict_treemap.html251 // live) and MapLibre (setScrollZoom toggles the GL handler), the robust fix is a capture-phaseCOMMENT
MEDIUMexamples/viz/smart_us_choropleth.html283 // live) and MapLibre (setScrollZoom toggles the GL handler), the robust fix is a capture-phaseCOMMENT
MEDIUMexamples/viz/smart_world_choropleth.html291 // live) and MapLibre (setScrollZoom toggles the GL handler), the robust fix is a capture-phaseCOMMENT
MEDIUMexamples/viz/smart_sales.html290 // live) and MapLibre (setScrollZoom toggles the GL handler), the robust fix is a capture-phaseCOMMENT
MEDIUMexamples/viz/smart_smarter.html274 // live) and MapLibre (setScrollZoom toggles the GL handler), the robust fix is a capture-phaseCOMMENT
MEDIUMexamples/viz/smart_dict_sunburst.html323 // live) and MapLibre (setScrollZoom toggles the GL handler), the robust fix is a capture-phaseCOMMENT
MEDIUMexamples/viz/smart_sales_kpi.html338 // live) and MapLibre (setScrollZoom toggles the GL handler), the robust fix is a capture-phaseCOMMENT
MEDIUMexamples/viz/smart_timeseries.html290 // live) and MapLibre (setScrollZoom toggles the GL handler), the robust fix is a capture-phaseCOMMENT
MEDIUMexamples/viz/smart_geo_outliers.html323 // live) and MapLibre (setScrollZoom toggles the GL handler), the robust fix is a capture-phaseCOMMENT
MEDIUMexamples/viz/gen_gallery.py354 # scrolls the page. The robust fix is a capture-phase wheel interceptor: inline we stopPropagationCOMMENT
MEDIUMbenchmarks/index_superpowers.html184 // live) and MapLibre (setScrollZoom toggles the GL handler), the robust fix is a capture-phaseCOMMENT
MEDIUMbenchmarks/stats_growth.html184 // live) and MapLibre (setScrollZoom toggles the GL handler), the robust fix is a capture-phaseCOMMENT
MEDIUMbenchmarks/trend.html184 // live) and MapLibre (setScrollZoom toggles the GL handler), the robust fix is a capture-phaseCOMMENT
MEDIUMbenchmarks/validate_growth.html184 // live) and MapLibre (setScrollZoom toggles the GL handler), the robust fix is a capture-phaseCOMMENT
MEDIUMbenchmarks/time_spent.html184 // live) and MapLibre (setScrollZoom toggles the GL handler), the robust fix is a capture-phaseCOMMENT
MEDIUMbenchmarks/sqlp_tuning.html184 // live) and MapLibre (setScrollZoom toggles the GL handler), the robust fix is a capture-phaseCOMMENT
MEDIUMbenchmarks/freq_growth.html184 // live) and MapLibre (setScrollZoom toggles the GL handler), the robust fix is a capture-phaseCOMMENT
MEDIUMbenchmarks/gainers.html184 // live) and MapLibre (setScrollZoom toggles the GL handler), the robust fix is a capture-phaseCOMMENT
MEDIUMbenchmarks/heatmap.html184 // live) and MapLibre (setScrollZoom toggles the GL handler), the robust fix is a capture-phaseCOMMENT
MEDIUMbenchmarks/index_advantage.html184 // live) and MapLibre (setScrollZoom toggles the GL handler), the robust fix is a capture-phaseCOMMENT
MEDIUMbenchmarks/count_callout.html184 // live) and MapLibre (setScrollZoom toggles the GL handler), the robust fix is a capture-phaseCOMMENT
MEDIUMbenchmarks/change_by_family.html184 // live) and MapLibre (setScrollZoom toggles the GL handler), the robust fix is a capture-phaseCOMMENT
MEDIUMscripts/pgo-train.sh3# pgo-train.sh - lean, single-pass PGO training harness for qsv.COMMENT
MEDIUMscripts/pgo-train.sh21# (see the `t` helper), so the same harness works against full and reduced builds.COMMENT
MEDIUMsrc/help_markdown_gen.rs7// Uses qsv-docopt Parser for robust structured parsing of options, arguments and defaults,COMMENT
MEDIUMsrc/mcp_skills_gen.rs7// Uses qsv-docopt Parser for robust USAGE text parsing.COMMENT
MEDIUMsrc/mcp_skills_gen.rs269 /// Parse USAGE text using qsv-docopt Parser for robust parsingCOMMENT
MEDIUMsrc/cmd/geocode.rs3830/// `US.NY.119` -> `119`), which is robust regardless of its leading digit.COMMENT
MEDIUMsrc/cmd/clean.rs335 // extension) as a sibling of the cache — robust to cwd, symlinks andCOMMENT
MEDIUMsrc/cmd/count.rs588 // now leverage the magic of Polars SQL with its lazy evaluation, to count the recordsCOMMENT
MEDIUMsrc/cmd/validate.rs1297 // This is robust to whitespace and avoids false matches onCOMMENT
MEDIUMsrc/cmd/pivotp.rs711 // moarstats: robust vs non-robust spread divergence — outliers inflate stddevCOMMENT
MEDIUMsrc/cmd/pivotp.rs737 // moarstats: robust quartile-based variabilityCOMMENT
MEDIUMsrc/cmd/moarstats.rs956/// Tukey's trimean - a robust estimator of central tendency thatCOMMENT
MEDIUMsrc/cmd/moarstats.rs968/// Midpoint of the middle 50% of data, a robust central tendency measure.COMMENT
MEDIUMsrc/cmd/moarstats.rs979/// Uses robust measures (MAD and median magnitude) instead of stddev and mean.COMMENT
MEDIUMsrc/cmd/moarstats.rs3863 // Trimean: (Q1 + 2*median + Q3) / 4 - Tukey's robust central tendency estimatorCOMMENT
MEDIUMsrc/cmd/moarstats.rs4123 // Add robust spread ratios (replace "mean" with empty string and clean up doubleCOMMENT
Structural Annotation Overuse53 hits · 101 pts
SeverityFileLineSnippetContext
LOWresources/describegpt_defaults.toml15### NOTE: The following VARIABLES are available in MiniJinja templates (use {{ variable }} syntax):COMMENT
LOWresources/describegpt_md_defaults.toml6### NOTE: The following VARIABLES are available in MiniJinja templates (use {{ variable }} syntax):COMMENT
LOWtests/test_sqlp.rs4573 // Step 1: Generate stats cache with --infer-dates so DateTime type is detectedCOMMENT
LOWtests/test_sqlp.rs4588 // Step 2: Generate Polars schema from stats cacheCOMMENT
LOWtests/test_sqlp.rs4605 // Step 3: Use sqlp with the cached schema to verify it parses datetimes correctlyCOMMENT
LOWtests/test_to.rs1449 // Step 1: Run `qsv schema --polars data.csv` to generate pschemaCOMMENT
LOWtests/test_to.rs1461 // Step 2: Run `qsv to parquet out/ data.csv` — should pick up the pschemaCOMMENT
LOW.claude/skills/review-respond/SKILL.md11## Step 1: Identify the PRCOMMENT
LOW.claude/skills/review-respond/SKILL.md19## Step 2: Fetch all pending review commentsCOMMENT
LOW.claude/skills/review-respond/SKILL.md31## Step 3: Route comments by reviewer typeCOMMENT
LOW.claude/skills/review-respond/SKILL.md36## Step 4: Apply fixes for human reviewer commentsCOMMENT
LOW.claude/skills/review-respond/SKILL.md52## Step 5: Verification gateCOMMENT
LOW.claude/skills/review-respond/SKILL.md62## Step 6: Run testsCOMMENT
LOW.claude/skills/review-respond/SKILL.md70## Step 7: CommitCOMMENT
LOW.claude/skills/review-respond/SKILL.md77## Step 8: Reply to resolved commentsCOMMENT
LOW.claude/skills/review-respond/SKILL.md85## Step 9: Report summaryCOMMENT
LOW.claude/skills/tests/duckdb.test.ts575 // Step 1: Convert CSV to Parquet using qsv_to_parquet (the MCP tool)COMMENT
LOW.claude/skills/tests/duckdb.test.ts586 // Step 2: Query the Parquet file with DuckDBCOMMENT
LOW.claude/skills/tests/pipeline-manifest.test.ts349 // Step 1: input.csv → intermediate.csvCOMMENT
LOW.claude/skills/tests/pipeline-manifest.test.ts383 // Step 1: read data.csv as inputCOMMENT
LOW.claude/skills/tests/pipeline-manifest.test.ts391 // Step 2: overwrite data.csv as output (e.g., in-place sort)COMMENT
LOW.claude/skills/tests/pipeline-manifest.test.ts793 // Step 1: reads data.csv (input), step 2: writes data.csv (output) — in-place overwriteCOMMENT
LOW.claude/skills/docs/desktop/README-MCPB.md25### Step 1: DownloadCOMMENT
LOW.claude/skills/docs/desktop/README-MCPB.md31### Step 2: Install in Claude DesktopCOMMENT
LOW.claude/skills/docs/desktop/README-MCPB.md35### Step 3: There is no step 3! The extension is ready to use. Just restart Claude Desktop and start asking Claude aboutCOMMENT
LOW.claude/skills/docs/guides/CLAUDE_CODE.md62### Step 1: Build the MCP ServerCOMMENT
LOW.claude/skills/docs/guides/CLAUDE_CODE.md73### Step 2: Configure Claude CodeCOMMENT
LOW.claude/skills/docs/guides/CLAUDE_CODE.md123### Step 3: Restart Claude CodeCOMMENT
LOW.claude/skills/docs/guides/START_HERE.md24## Step 1: Install the MCP ServerCOMMENT
LOW.claude/skills/docs/guides/GEMINI_CLI.md25### Step 1: Build the MCP ServerCOMMENT
LOW.claude/skills/docs/guides/GEMINI_CLI.md37### Step 2: Configure Gemini CLICOMMENT
LOW.claude/skills/scripts/install-mcp.js344 // Step 1: Check for qsv binaryCOMMENT
LOW.claude/skills/scripts/install-mcp.js350 // Step 2: Build TypeScriptCOMMENT
LOW.claude/skills/scripts/install-mcp.js356 // Step 3: Configure environment variablesCOMMENT
LOW.claude/skills/scripts/install-mcp.js360 // Step 4: Update Claude Desktop configCOMMENT
LOW.claude/skills/scripts/install-mcp.js366 // Step 5: Print verification stepsCOMMENT
LOW.claude/skills/scripts/package-mcpb.js271 // Step 1: Clean upCOMMENT
LOW.claude/skills/scripts/package-mcpb.js274 // Step 2: Build TypeScriptCOMMENT
LOW.claude/skills/scripts/package-mcpb.js277 // Step 3: Validate filesCOMMENT
LOW.claude/skills/scripts/package-mcpb.js280 // Step 4: Create archiveCOMMENT
LOW.claude/skills/scripts/package-mcpb.js283 // Step 5: Display summaryCOMMENT
LOW.claude/skills/skills/bls-query/SKILL.md15### Step 1: Identify the topicCOMMENT
LOW.claude/skills/skills/bls-query/SKILL.md29### Step 2: Choose the right toolCOMMENT
LOW.claude/skills/skills/bls-query/SKILL.md39### Step 3: Interpret and present resultsCOMMENT
LOW.claude/skills/skills/bls-query/SKILL.md47### Step 4: Handle unknownsCOMMENT
LOW.claude/skills/src/tool-handlers.ts1498 // Step 1: Ensure stats cache is up-to-date (needed by both DuckDB and qsv to parquet paths)COMMENT
LOW.claude/skills/src/tool-handlers.ts1501 // Step 3: Convert to Parquet (DuckDB with ZSTD when available, qsv to parquet with ZSTD otherwise)COMMENT
LOW.claude/skills/src/parquet-bridge.ts689 // Step 2: Generate Polars schemaCOMMENT
LOW.claude/skills/src/parquet-bridge.ts718 // Step 3: Convert to Parquet (DuckDB with ZSTD when available, qsv to parquet with ZSTD otherwise)COMMENT
LOWdocs/contributor/STATS_TECHNICAL_GUIDE.md738 // Step 1: Parse command-line argumentsCOMMENT
LOWdocs/contributor/STATS_TECHNICAL_GUIDE.md741 // Step 2: Handle typesonly mode (disable other stats)COMMENT
LOWdocs/contributor/STATS_TECHNICAL_GUIDE.md748 // Step 3: Setup boolean inferenceCOMMENT
LOWdocs/contributor/STATS_TECHNICAL_GUIDE.md754 // Step 4: Check environment variable overridesCOMMENT
Fake / Example Data66 hits · 56 pts
SeverityFileLineSnippetContext
LOWtests/test_fixedwidth.rs89 svec!["Jane Doe", "1017"],CODE
LOWtests/test_geocode.rs2016 .args(["--api-key", "dummy-key-for-arg-validation"])CODE
LOWtests/test_geocode.rs2132 .args(["--api-key", "dummy-key-for-arg-validation"])CODE
LOWtests/test_geocode.rs2154 .args(["--api-key", "dummy-key-for-arg-validation"])CODE
LOWtests/test_geocode.rs2178 .args(["--api-key", "dummy-key-for-arg-validation"])CODE
LOWtests/test_geocode.rs2196 .args(["--api-key", "dummy-key-for-arg-validation"])CODE
LOWtests/test_geocode.rs2214 .args(["--api-key", "dummy-key-for-arg-validation"])CODE
LOWtests/test_geocode.rs2231 .args(["--api-key", "dummy-key-for-arg-validation"])CODE
LOWtests/test_profile.rs1813 "contact_point": {"fn": "Jane Doe", "hasEmail": "jane@example.gov"},CODE
LOWtests/test_profile.rs2135 "contact_point": {"fn": "Jane Doe", "hasEmail": "jane@example.gov"},CODE
LOWtests/test_validate.rs3034 "user@example.com",CODE
LOWtests/test_validate.rs3041 "user@example.com",CODE
LOWtests/test_validate.rs3048 "user@example.com",CODE
LOWtests/test_validate.rs3139 "user@example.com",CODE
LOWtests/test_validate.rs3146 "user@example.com",CODE
LOWtests/test_validate.rs3153 "user@example.com",CODE
LOWtests/test_validate.rs758 svec!["1", "John Doe", "john@example.com", "IT"],CODE
LOWtests/test_validate.rs760 svec!["3", "John Doe", "john@example.com", "IT"], // Duplicate name+emailCODE
LOWtests/test_validate.rs802 svec!["1", "John Doe", "john@example.com", "IT"],CODE
LOWtests/test_validate.rs811 svec!["3", "John Doe", "john@example.com", "IT"],CODE
LOWtests/test_validate.rs828 svec!["1", "John Doe", "john@example.com", "IT"],CODE
LOWtests/test_validate.rs830 svec!["3", "John Doe", "john@example.com", "IT"], // Duplicate name+emailCODE
LOWtests/test_validate.rs872 svec!["1", "John Doe", "john@example.com", "IT"],CODE
LOWtests/test_validate.rs881 svec!["3", "John Doe", "john@example.com", "IT"],CODE
LOWtests/test_validate.rs898 svec!["1", "John Doe", "john@example.com", "IT"],CODE
LOWtests/test_validate.rs900 svec!["3", "John Doe", "john@example.com", "IT"], // Duplicate name+emailCODE
LOWtests/test_validate.rs942 svec!["1", "John Doe", "john@example.com", "IT"],CODE
LOWtests/test_validate.rs951 svec!["3", "John Doe", "john@example.com", "IT"],CODE
LOWtests/test_validate.rs968 svec!["1", "John Doe", "john@example.com", "IT"],CODE
LOWtests/test_validate.rs970 svec!["3", "John Doe", "", "IT"], // Empty emailCODE
LOWtests/test_validate.rs1021 svec!["1", "John Doe", "john@example.com", "IT"],CODE
LOWtests/test_validate.rs1023 svec!["3", "John Doe", "", "IT"],CODE
LOWtests/test_validate.rs1045 svec!["1", "John Doe", "john.doe@example.com", "IT"],CODE
LOWtests/test_validate.rs1047 svec!["3", "John Doe", "john.doe@example.com", "IT"], // DuplicateCODE
LOWtests/test_validate.rs1089 svec!["1", "John Doe", "john.doe@example.com", "IT"],CODE
LOWtests/test_validate.rs1098 svec!["3", "John Doe", "john.doe@example.com", "IT"],CODE
LOWtests/test_validate.rs1316 svec!["1", "John Doe", "john@example.com", "IT", "Developer"],CODE
LOWtests/test_validate.rs1318 svec!["3", "John Doe", "john@example.com", "IT", "Developer"], // Duplicate name+email+roleCODE
LOWtests/test_validate.rs1367 svec!["1", "John Doe", "john@example.com", "IT", "Developer"],CODE
LOWtests/test_validate.rs1378 svec!["3", "John Doe", "john@example.com", "IT", "Developer"],CODE
LOWtests/test_validate.rs1397 "John Doe",CODE
LOWtests/test_validate.rs1835 svec!["1", "John Doe", "not-an-email", "not-a-url", "not-currency"],CODE
LOWtests/test_validate.rs1919 svec!["1", "John Doe", "not-an-email", "25"], // Format error onlyCODE
LOWtests/test_validate.rs1986 svec!["1", "John Doe", "not-an-email", "25"], // Now valid (format ignored)CODE
LOWtests/test_validate.rs2615 svec!["1", "John Doe", "user@example.com"], // Valid emailCODE
LOWtests/test_validate.rs2615 svec!["1", "John Doe", "user@example.com"], // Valid emailCODE
LOWtests/test_validate.rs2661 svec!["1", "John Doe", "user@example.com"],CODE
LOWtests/test_validate.rs2661 svec!["1", "John Doe", "user@example.com"],CODE
LOWtests/test_validate.rs2688 svec!["1", "John Doe", "user@example.com"],CODE
LOWtests/test_validate.rs2688 svec!["1", "John Doe", "user@example.com"],CODE
LOWtests/test_validate.rs2741 svec!["1", "John Doe", "user@example.com"],CODE
LOWtests/test_validate.rs2741 svec!["1", "John Doe", "user@example.com"],CODE
LOWtests/test_validate.rs2796 svec!["1", "John Doe", "user@example.com"],CODE
LOWtests/test_validate.rs2796 svec!["1", "John Doe", "user@example.com"],CODE
LOWtests/test_validate.rs3020 "user@example.com",CODE
LOWtests/test_validate.rs3112 "user@example.com",CODE
LOWtests/test_luau.rs3409 svec!["user@example.com", "yes"],CODE
LOWtests/test_luau.rs3422 let expected = vec![svec!["email", "valid"], svec!["user@example.com", "yes"]];CODE
LOWtests/resources/profile/dcat-init-context.json17 "contact_point": {"fn": "Jane Doe", "hasEmail": "jane.doe@example.gov"},CODE
LOW…/profile/golden/wprdc-311-subset.dataset.expected.json15 "vcard:fn": "Jane Doe",CODE
6 more matches not shown…
Verbosity Indicators25 hits · 48 pts
SeverityFileLineSnippetContext
LOWtests/test_sqlp.rs4573 // Step 1: Generate stats cache with --infer-dates so DateTime type is detectedCOMMENT
LOWtests/test_sqlp.rs4588 // Step 2: Generate Polars schema from stats cacheCOMMENT
LOWtests/test_sqlp.rs4605 // Step 3: Use sqlp with the cached schema to verify it parses datetimes correctlyCOMMENT
LOWtests/test_to.rs1449 // Step 1: Run `qsv schema --polars data.csv` to generate pschemaCOMMENT
LOWtests/test_to.rs1461 // Step 2: Run `qsv to parquet out/ data.csv` — should pick up the pschemaCOMMENT
LOW.claude/skills/tests/duckdb.test.ts575 // Step 1: Convert CSV to Parquet using qsv_to_parquet (the MCP tool)COMMENT
LOW.claude/skills/tests/duckdb.test.ts586 // Step 2: Query the Parquet file with DuckDBCOMMENT
LOW.claude/skills/tests/pipeline-manifest.test.ts349 // Step 1: input.csv → intermediate.csvCOMMENT
LOW.claude/skills/tests/pipeline-manifest.test.ts383 // Step 1: read data.csv as inputCOMMENT
LOW.claude/skills/tests/pipeline-manifest.test.ts391 // Step 2: overwrite data.csv as output (e.g., in-place sort)COMMENT
LOW.claude/skills/tests/pipeline-manifest.test.ts793 // Step 1: reads data.csv (input), step 2: writes data.csv (output) — in-place overwriteCOMMENT
LOW.claude/skills/scripts/install-mcp.js344 // Step 1: Check for qsv binaryCOMMENT
LOW.claude/skills/scripts/install-mcp.js350 // Step 2: Build TypeScriptCOMMENT
LOW.claude/skills/scripts/install-mcp.js356 // Step 3: Configure environment variablesCOMMENT
LOW.claude/skills/scripts/install-mcp.js360 // Step 4: Update Claude Desktop configCOMMENT
LOW.claude/skills/scripts/install-mcp.js366 // Step 5: Print verification stepsCOMMENT
LOW.claude/skills/scripts/package-mcpb.js271 // Step 1: Clean upCOMMENT
LOW.claude/skills/scripts/package-mcpb.js274 // Step 2: Build TypeScriptCOMMENT
LOW.claude/skills/scripts/package-mcpb.js277 // Step 3: Validate filesCOMMENT
LOW.claude/skills/scripts/package-mcpb.js280 // Step 4: Create archiveCOMMENT
LOW.claude/skills/scripts/package-mcpb.js283 // Step 5: Display summaryCOMMENT
LOW.claude/skills/src/tool-handlers.ts1498 // Step 1: Ensure stats cache is up-to-date (needed by both DuckDB and qsv to parquet paths)COMMENT
LOW.claude/skills/src/tool-handlers.ts1501 // Step 3: Convert to Parquet (DuckDB with ZSTD when available, qsv to parquet with ZSTD otherwise)COMMENT
LOW.claude/skills/src/parquet-bridge.ts689 // Step 2: Generate Polars schemaCOMMENT
LOW.claude/skills/src/parquet-bridge.ts718 // Step 3: Convert to Parquet (DuckDB with ZSTD when available, qsv to parquet with ZSTD otherwise)COMMENT
Cross-Language Confusion10 hits · 45 pts
SeverityFileLineSnippetContext
HIGHexamples/viz/gen_gallery.py181 "<script>(function(){var target=null,until=0,iv=null,moved=false;"CODE
HIGHexamples/viz/gen_gallery.py182 "function stop(){if(iv){clearInterval(iv);iv=null;}}"CODE
HIGHexamples/viz/gen_gallery.py222 "b._t=setTimeout(function(){b.classList.remove(\"ok\");b.title=\"Copy\";b._t=null;},1200);}"CODE
HIGHexamples/viz/gen_gallery.py308 'if(!b||typeof b.z!=="number"||!(plotW>0)||!(plotH>0))return null;'CODE
HIGHexamples/viz/gen_gallery.py312 'var ratio=Math.min((dx*plotW)/aw,(dy*plotH)/ah);if(!isFinite(ratio)||ratio<=0)return null;'CODE
HIGHexamples/viz/gen_gallery.py326 'function qsvFitNow(gd,tries){if(tries===undefined)tries=20;qsvCaptureBaked(gd);'CODE
HIGHexamples/viz/gen_gallery.py346 'function qsvApplyScrollZoom(gd,tries){if(tries===undefined)tries=20;var fl=gd._fullLayout||{};'CODE
HIGHexamples/viz/gen_gallery.py396 " cargo build --bin qsv -F all_features && python3 examples/viz/gen_gallery.py\n"CODE
HIGHscripts/gen_benchmark_viz.py19 git add benchmarks && git commit # Pages redeploys on push to masterSTRING
HIGHscripts/gen_benchmark_viz.py33 git -C qsv.wiki add Benchmarks.md && git -C qsv.wiki commit && git -C qsv.wiki pushSTRING
Modern AI Meta-Vocabulary8 hits · 21 pts
SeverityFileLineSnippetContext
MEDIUMCHANGELOG.md1145- **Seamlessly works with both [Claude Code](https://code.claude.com/docs/en/overview) and the just launched [Claude CowCODE
MEDIUMresources/describegpt_md_defaults.toml13# {{ reasoning }} - LLM reasoning trace (may be empty)COMMENT
MEDIUMdocs/Describegpt.md231## SQL Query Generation and Execution ("SQL RAG" mode)COMMENT
MEDIUMdocs/help/TableOfContents.md19| [describegpt](describegpt.md)<br>[📇](#legend "uses an index when available.")[🗃️](#legend "Limited Extended input suppCODE
MEDIUMdocs/help/TableOfContents.md19| [describegpt](describegpt.md)<br>[📇](#legend "uses an index when available.")[🗃️](#legend "Limited Extended input suppCODE
MEDIUMexamples/viz/gen_gallery.py993 # reuse the existing scaffold verbatim: everything up to and including `<div class="grid">`,COMMENT
MEDIUMsrc/cmd/describegpt.rs88 # Ask detailed natural language questions that require SQL queries and auto-invoke SQL RAG modeCOMMENT
MEDIUMsrc/cmd/describegpt.rs3595 // Schema scaffold. Cache that shape regardless of the chosen output formatCOMMENT
Slop Phrases6 hits · 16 pts
SeverityFileLineSnippetContext
MEDIUM.serena/project.yml44# Same syntax as gitignore, so you can use * and **.COMMENT
MEDIUM.claude/skills/.serena/project.yml44# Same syntax as gitignore, so you can use * and **.COMMENT
MEDIUMsrc/cmd/sqlp.rs72 # In long, complex scripts that produce multiple temporary tables, note that you can useCOMMENT
MEDIUMsrc/cmd/sqlp.rs136 # note that you can also use read_csv() to read compressed files directlyCOMMENT
MEDIUMsrc/cmd/sqlp.rs142 # apart from using Polar's table functions, you can also use SKIP_INPUT when the SELECTCOMMENT
LOWsrc/cmd/replace.rs5the replacement string. But don't forget to escape your $ in bash by using aCODE
Cross-File Repetition3 hits · 15 pts
SeverityFileLineSnippetContext
HIGHtests/test_py.rs0{qty} {fruit} cost ${(float(unit_cost) * float(qty)):.2f}. its quite {"cheap" if ((float(unit_cost) * float(qty)) < 20.0STRING
HIGHdocs/help/py.md0{qty} {fruit} cost ${(float(unit_cost) * float(qty)):.2f}. its quite {"cheap" if ((float(unit_cost) * float(qty)) < 20.0STRING
HIGHsrc/cmd/python.rs0{qty} {fruit} cost ${(float(unit_cost) * float(qty)):.2f}. its quite {"cheap" if ((float(unit_cost) * float(qty)) < 20.0STRING
Hyper-Verbose Identifiers14 hits · 14 pts
SeverityFileLineSnippetContext
LOW.claude/skills/tests/elicitation.test.ts72async function buildDirectorySuggestions(currentWorkingDir: string): Promise<string> {CODE
LOW.claude/skills/src/file-operations.ts261export async function resolveAndConvertInputFile(CODE
LOW.claude/skills/src/file-operations.ts623export function collectAdditionalInputFiles(CODE
LOW.claude/skills/src/command-guidance.ts188export function enhanceParameterDescription(CODE
LOW.claude/skills/src/tool-handlers.ts403async function runSqlpParquetInterception(CODE
LOW.claude/skills/src/tool-handlers.ts493export function validateLlmResponsesShape(arr: unknown[]): string | null {CODE
LOW.claude/skills/src/tool-handlers.ts530async function runDescribegptInterception(CODE
LOW.claude/skills/src/utils.ts196export function describegptFallbackResult(args: string[]): string {CODE
LOW.claude/skills/src/tool-definitions.ts108export function createGenericToolDefinition(CODE
LOW.claude/skills/src/tool-definitions.ts257export function createBrowseDirectoryTool(): McpToolDefinition {CODE
LOW.claude/skills/src/tool-definitions.ts444export function createSetupToolDefinition(): McpToolDefinition {CODE
LOW.claude/skills/src/version.ts76export function readMinimumQsvVersionFromManifest(projectRoot: string): string | null {CODE
LOW.claude/skills/src/mcp-sampling.ts119export async function executeDescribegptWithSampling(CODE
LOWscripts/docs-drift-check.py117def expand_all_features_for_docs(features: dict[str, set[str]]) -> set[str]:CODE
Self-Referential Comments4 hits · 12 pts
SeverityFileLineSnippetContext
MEDIUMsrc/cmd/applydp.rs115 # Create a new column 'mailing address' from 'house number', 'street', 'city' and 'zip-code' columns:COMMENT
MEDIUMsrc/cmd/applydp.rs118 # Create a new column 'FullName' from 'FirstName', 'MI', and 'LastName' columns:COMMENT
MEDIUMsrc/cmd/apply.rs186 # Create a new column 'mailing address' from 'house number', 'street', 'city'COMMENT
MEDIUMsrc/cmd/apply.rs190 # Create a new column 'FullName' from 'FirstName', 'MI', and 'LastName' columns:COMMENT
Deep Nesting9 hits · 9 pts
SeverityFileLineSnippetContext
LOW…laude/skills/visual-data-dictionary/edit_dictionary.py274CODE
LOW…laude/skills/visual-data-dictionary/edit_dictionary.py363CODE
LOWdocs/describegpt/check_semanticmd.py73CODE
LOWdocs/describegpt/check_semanticmd.py118CODE
LOWexamples/viz/gen_gallery.py831CODE
LOWexamples/viz/gen_gallery.py983CODE
LOWexamples/viz/gen_world_cities.py48CODE
LOWscripts/docs-drift-check.py196CODE
LOWscripts/gen_benchmark_viz.py192CODE
Synthetic Comment Markers1 hit · 8 pts
SeverityFileLineSnippetContext
HIGHCHANGELOG.md1682* Created [Event Logo Archive](https://github.com/dathere/qsv/tree/master/docs/images/event-logos) with AI-generated seaCOMMENT
Redundant / Tautological Comments4 hits · 6 pts
SeverityFileLineSnippetContext
LOW.serena/project.yml79# Set this to [] to disable base modes for this project.COMMENT
LOW.serena/project.yml80# Set this to a list of mode names to always include the respective modes for this project.COMMENT
LOWscripts/benchmarks.sh902# Check if a results directory exists, if it doesn't create itCOMMENT
LOWsrc/cmd/sortcheck.rs35 # Check if file.csv is lexicographically sorted on all columns:COMMENT
Unused Imports2 hits · 2 pts
SeverityFileLineSnippetContext
LOWdocs/describegpt/check_semanticmd.py27CODE
LOWscripts/docs-drift-check.py34CODE
Example Usage Blocks1 hit · 2 pts
SeverityFileLineSnippetContext
LOWscripts/qsv-tune.sh29# Usage:COMMENT
Excessive Try-Catch Wrapping1 hit · 1 pts
SeverityFileLineSnippetContext
LOW…laude/skills/visual-data-dictionary/edit_dictionary.py244 except Exception:CODE