Repository Analysis

spiceai/spiceai

Add a real-time analytics node to your operational database. Spice is a portable, accelerated SQL query, search, and LLM-inference engine in Rust for data-grounded AI apps and agents.

4.4 Likely human-written View on GitHub

Analysis Overview

This report presents the forensic synthetic code analysis of spiceai/spiceai, a Rust project with 3,073 GitHub stars. SynthScan v2.0 examined 1,363,258 lines of code across 4139 source files, recording 4723 pattern matches distributed across 17 syntactic categories. The overall adjusted score of 4.4 places this repository in the Likely human-written band.

The scanner applied 160+ deterministic lexical heuristics, multi-line block detectors, abstract syntax tree depth profilers, and a cross-file Jaccard similarity matrix to construct a statistically normalised synthetic code estimate. All matches are individually weighted by severity coefficient and contextual multiplier before summation, and the resulting headline score is temporally discounted to account for the repository's development history relative to the commercial emergence of large language model coding tooling (November 2022 onward).

4.4
Adjusted Score
4.4
Raw Score
100%
Time Factor
2026-08-29
Last Push
3.1K
Stars
Rust
Language
1.4M
Lines of Code
4.1K
Files
4.7K
Pattern Hits
2026-08-29
Scan Date
0.00
HC Hit Rate

What These Metrics Mean

Adjusted Score
Primary synthetic code indicator. Raw score normalised per 1,000 lines of code and multiplied by the temporal discount factor. This is the definitive comparative metric — use it to rank repositories by AI authorship density.
Raw Score
The unmodified sum of all severity-weighted, context-multiplied pattern match scores before temporal discounting. Reflects the absolute signal strength independent of when the repository was last active.
Time Factor
The temporal discount multiplier (0–100%) applied to the raw score. Repositories last updated before ChatGPT's launch (Nov 2022) receive a 5% factor. Full signal is only assigned to repositories active in the post-adoption era (Jan 2024+).
Pattern Hits
Total count of individual pattern matches across all files and categories. A high hit count with a low score may indicate a very large codebase with isolated AI snippets; a low count with a high score indicates dense, concentrated AI signatures.
HC Hit Rate
High+Critical pattern hits per file, averaged across the repository. This orthogonal signal catches repositories where a few files are densely packed with high-severity AI tells — a strong indicator even when the normalised score appears moderate due to codebase size.
Lines of Code / Files
Total lines and files analysed. The scanner examines 94 file extensions. These denominators are used to normalise the score, enabling fair comparison between repositories of vastly different sizes.

Score History

This chart maps the temporal evolution of the adjusted synthetic code score across successive scan runs. An upward trajectory indicates ongoing incorporation of AI-generated code or expanding LLM-assisted scaffolding; a stable or declining trajectory may reflect active human refactoring, code removal, or the adoption of stricter authorship policies. The dashed secondary line (right axis) independently tracks total raw pattern hit count, which can diverge from the normalised score when codebase size changes significantly between scans.

Severity Breakdown

Classifies detected patterns by their diagnostic confidence and structural impact. CRITICAL patterns (coefficient 10) represent definitive synthetic signatures — hallucinated imports, explicit LLM attribution metadata — virtually never produced by human authors. HIGH (5) indicates strong structural tells such as cross-file repetition or cross-linguistic idioms. MEDIUM (2) covers recognisable conversational padding and AI-specific vocabulary. LOW (1) captures subtle indicators like tautological comments and generic boilerplate that require density to carry independent signal.

CRITICAL 0HIGH 4MEDIUM 626LOW 4093

Directory Score Breakdown

This horizontal bar chart decomposes the repository's raw synthetic code score by top-level directory, allowing you to pinpoint precisely which modules or components carry the highest AI authorship density. Directories with disproportionately high scores relative to their size warrant targeted manual review: concentrated AI signatures often trace back to mass-generated configuration layers, auto-ported test suites, LLM-scaffolded boilerplate classes, or entire subsystems authored under heavy copilot assistance. Use this view to prioritise your human code-review effort.

Pattern Findings

The scanner identified 4723 distinct pattern matches across 17 syntactic categories. Each entry below represents a discrete location in the source code where the engine recorded a statistically significant AI authorship indicator. Expand any category row to inspect the individual file paths, line numbers, code snippets, and the lexical context (CODE, COMMENT, or STRING) in which each match was detected.

Reading the findings table: The Severity column indicates the diagnostic confidence level (CRITICAL / HIGH / MEDIUM / LOW). The Context column identifies whether the match occurred inside executable code, an inline comment, or a string literal — comment-context matches receive a ×1.5 weight because LLMs systematically over-annotate. The ⚡ bolt icon marks clustered matches: three or more patterns within a 10-line window, each receiving an additional ×1.5 density multiplier as dense clusters constitute far stronger evidence of synthetic authorship than isolated hits.

Over-Commented Block3509 hits · 3408 pts
SeverityFileLineSnippetContext
LOWCargo.toml141# crate that forgets to opt in. NOTE: `-Dwarnings` intentionally stays in theCOMMENT
LOWdeny.toml21 # Local model search dependencies pull `fxhash`; no safe upgrade is available yet. Owner: llms.COMMENT
LOWlayers.toml1# Crate layering manifest — see docs/dev/crate_layering.mdCOMMENT
LOWlayers.toml21# building blocks that no runtime code touches live here.COMMENT
LOWlayers.toml61# orchestrator (they may still use the low runtime-*-api interface crates).COMMENT
LOWlayers.toml81# deps are exempt. Ratchet: add one entry per source as its driver is fully evacuatedCOMMENT
LOWlayers.toml161# interface crates a connector needs (data-connector-api) name a component'sCOMMENT
LOWlayers.toml181# move up once that connector is evacuated (see docs/dev/crate_layering.md).COMMENT
LOWlayers.toml201# `runtime::token_providers::databricks`; it flips with the catalog seam.COMMENT
LOWlayers.toml301COMMENT
LOW.config/nextest.toml1[profile.default]COMMENT
LOW.config/nextest.toml21# the reference model on every run. These are deterministic correctness checks,COMMENT
LOW.config/nextest.toml41# mutation_property mixed_dense_position 122.8s 208.6s 1.70xCOMMENT
LOW.config/nextest.toml61# that is a regression to investigate, not a reason to raise this again. It isCOMMENT
LOW.config/nextest.toml81# group above means these tests are expected to pass on the first try under theCOMMENT
LOW.config/nextest.toml101# because the ceiling is already above what the model predicts under the worstCOMMENT
LOW…estoperator/dispatch/perf-cayenne-vs-duckdb/pairs.yaml1# Cayenne vs DuckDB performance comparison matrix.COMMENT
LOW…estoperator/dispatch/perf-cayenne-vs-duckdb/pairs.yaml21# scale_factor: 1 | 5 | 10 | 100 | 1000COMMENT
LOWtools/testoperator/src/system_adapter.rs21//! (the same protocol spicebench uses). The adapter's `setup()` response carriesCOMMENT
LOWtools/testoperator/src/system_adapter.rs81COMMENT
LOWtools/testoperator/src/pg_stats.rs41pub struct PgStatSample {COMMENT
LOWtools/testoperator/src/probe.rs21//! **Why the generator needs them.** From outside the process, a run that isCOMMENT
LOWtools/testoperator/src/probe.rs41//!COMMENT
LOWtools/testoperator/src/probe.rs61/// because the harness that reads them runs *on* the generator host (over ssh)COMMENT
LOWtools/testoperator/src/probe.rs121 }COMMENT
LOWtools/testoperator/src/commands/mod.rs501/// * shared — every thread multiplexes one connection.COMMENT
LOWtools/testoperator/src/commands/bench/mod.rs321/// and in what order — follows partitioning and aggregation order rather than theCOMMENT
LOWtools/testoperator/src/commands/bench/mod.rs381/// partitions and the low bits differ per machine. Result snapshots for theseCOMMENT
LOWtools/testoperator/src/commands/streaming/mutations.rs21//! before reaching their final TPC-H state, while the rest are inserted directly.COMMENT
LOWtools/testoperator/src/commands/streaming/mutations.rs321 result.push(batch);COMMENT
LOW…ols/testoperator/src/commands/streaming/correctness.rs21//! seed to exercise different mutation patterns.COMMENT
LOW…ls/testoperator/src/commands/streaming/verification.rs41/// Result of running TPCH verification using `SpiceTest`.COMMENT
LOWtools/testoperator/src/commands/streaming/runner.rs21//! correctness with TPC-H queries.COMMENT
LOWtools/testoperator/src/commands/streaming/traits.rs61COMMENT
LOWtools/testoperator/src/commands/streaming/utils.rs241/// `DynamoDB` metrics fetched from Spice's Prometheus endpoint.COMMENT
LOW…estoperator/src/commands/streaming/sources/dynamodb.rs1161 async fn cleanup(&self) -> Result<()> {COMMENT
LOWtools/testoperator/src/commands/search/mteb.rs21};COMMENT
LOWtools/testoperator/src/commands/htap/mod.rs81COMMENT
LOWtools/testoperator/src/commands/htap/reporting.rs41 /// Worst (across datasets) P99 lag over the under-load window, in milliseconds.COMMENT
LOWtools/testoperator/src/commands/htap/reporting.rs241 }COMMENT
LOWtools/testoperator/src/commands/htap/reporting.rs581#[expect(clippy::cast_possible_truncation, clippy::cast_sign_loss)]COMMENT
LOWtools/testoperator/src/commands/htap/reporting.rs961/// JSON — the durable, machine-readable artifact `scripts/chbench-waterfall.py`COMMENT
LOW…stoperator/src/commands/htap/correctness/analytical.rs41};COMMENT
LOW…stoperator/src/commands/htap/correctness/analytical.rs201 }COMMENT
LOW…stoperator/src/commands/htap/correctness/analytical.rs241 // `workers` `evaluate_query` futures run concurrently, and a slot is freed asCOMMENT
LOW…stoperator/src/commands/htap/correctness/analytical.rs341 outcome: Outcome::SpiceError(e.to_string()),COMMENT
LOW…/testoperator/src/commands/htap/correctness/compare.rs21//! the same logical values with different physical Arrow encodings. TheseCOMMENT
LOW…/testoperator/src/commands/htap/correctness/compare.rs41//! decision is made elsewhere.COMMENT
LOW…/testoperator/src/commands/htap/correctness/compare.rs61pub struct NumericDelta {COMMENT
LOW…/testoperator/src/commands/htap/correctness/compare.rs121/// order-dependent rounding), so the fingerprint gate can compare it with zeroCOMMENT
LOW…/testoperator/src/commands/htap/correctness/compare.rs141pub fn float_columns(batch: &RecordBatch) -> Vec<bool> {COMMENT
LOW…/testoperator/src/commands/htap/correctness/compare.rs201 /// A decimal pair that could not be brought to a common scale withoutCOMMENT
LOW…/testoperator/src/commands/htap/correctness/compare.rs261 let e = arrow::compute::cast(e_col, &DataType::Float64).ok()?;COMMENT
LOW…/testoperator/src/commands/htap/correctness/compare.rs521 /// Regression: the fingerprint gate compared exact columns via `f64` andCOMMENT
LOW…/testoperator/src/commands/htap/correctness/compare.rs701COMMENT
LOW…estoperator/src/commands/htap/correctness/row_count.rs41#[derive(Debug, Clone)]COMMENT
LOW…estoperator/src/commands/htap/correctness/row_count.rs81/// still reported as "replication did not converge").COMMENT
LOW…estoperator/src/commands/htap/correctness/row_count.rs101 }COMMENT
LOW…estoperator/src/commands/htap/correctness/row_count.rs461 }COMMENT
LOW…estoperator/src/commands/htap/correctness/row_count.rs581 audit.spice,COMMENT
3449 more matches not shown…
Decorative Section Separators551 hits · 1596 pts
SeverityFileLineSnippetContext
MEDIUM…estoperator/dispatch/perf-cayenne-vs-duckdb/pairs.yaml45 # ===========================================================================COMMENT
MEDIUM…estoperator/dispatch/perf-cayenne-vs-duckdb/pairs.yaml47 # ===========================================================================COMMENT
MEDIUM…estoperator/dispatch/perf-cayenne-vs-duckdb/pairs.yaml176 # ===========================================================================COMMENT
MEDIUM…estoperator/dispatch/perf-cayenne-vs-duckdb/pairs.yaml178 # ===========================================================================COMMENT
MEDIUM…estoperator/dispatch/perf-cayenne-vs-duckdb/pairs.yaml218 # ===========================================================================COMMENT
MEDIUM…estoperator/dispatch/perf-cayenne-vs-duckdb/pairs.yaml220 # ===========================================================================COMMENT
MEDIUM…estoperator/dispatch/perf-cayenne-vs-duckdb/pairs.yaml249 # ===========================================================================COMMENT
MEDIUM…estoperator/dispatch/perf-cayenne-vs-duckdb/pairs.yaml251 # ===========================================================================COMMENT
MEDIUM…estoperator/dispatch/perf-cayenne-vs-duckdb/pairs.yaml274 # ===========================================================================COMMENT
MEDIUM…estoperator/dispatch/perf-cayenne-vs-duckdb/pairs.yaml276 # ===========================================================================COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.secrets.yaml10# =============================================================================COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.secrets.yaml12# =============================================================================COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.catalogs.yaml2# =============================================================================COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.catalogs.yaml4# =============================================================================COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.catalogs.yaml10 # ---------------------------------------------------------------------------COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.catalogs.yaml13 # ---------------------------------------------------------------------------COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.catalogs.yaml22 # ---------------------------------------------------------------------------COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.catalogs.yaml24 # ---------------------------------------------------------------------------COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.catalogs.yaml31 # ---------------------------------------------------------------------------COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.catalogs.yaml33 # ---------------------------------------------------------------------------COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.catalogs.yaml40 # ---------------------------------------------------------------------------COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.catalogs.yaml42 # ---------------------------------------------------------------------------COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.catalogs.yaml48 # ---------------------------------------------------------------------------COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.catalogs.yaml50 # ---------------------------------------------------------------------------COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.embeddings.yaml2# =============================================================================COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.embeddings.yaml4# =============================================================================COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.embeddings.yaml10 # ---------------------------------------------------------------------------COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.embeddings.yaml12 # ---------------------------------------------------------------------------COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.embeddings.yaml23 # ---------------------------------------------------------------------------COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.embeddings.yaml26 # ---------------------------------------------------------------------------COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.embeddings.yaml35 # ---------------------------------------------------------------------------COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.embeddings.yaml37 # ---------------------------------------------------------------------------COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.embeddings.yaml43 # ---------------------------------------------------------------------------COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.embeddings.yaml45 # ---------------------------------------------------------------------------COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.runtime.yaml10# =============================================================================COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.runtime.yaml12# =============================================================================COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.datasets.yaml6# =============================================================================COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.datasets.yaml8# =============================================================================COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.datasets.yaml10 # ---------------------------------------------------------------------------COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.datasets.yaml13 # ---------------------------------------------------------------------------COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.datasets.yaml42 # ---------------------------------------------------------------------------COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.datasets.yaml44 # ---------------------------------------------------------------------------COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.datasets.yaml50 # ---------------------------------------------------------------------------COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.datasets.yaml55 # ---------------------------------------------------------------------------COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.datasets.yaml120 # ---------------------------------------------------------------------------COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.datasets.yaml123 # ---------------------------------------------------------------------------COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.datasets.yaml132 # ---------------------------------------------------------------------------COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.datasets.yaml135 # ---------------------------------------------------------------------------COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.datasets.yaml142 # ---------------------------------------------------------------------------COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.datasets.yaml145 # ---------------------------------------------------------------------------COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.datasets.yaml197 # ---------------------------------------------------------------------------COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.datasets.yaml200 # ---------------------------------------------------------------------------COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.datasets.yaml210 # ---------------------------------------------------------------------------COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.datasets.yaml213 # ---------------------------------------------------------------------------COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.datasets.yaml222 # ---------------------------------------------------------------------------COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.datasets.yaml229 # ---------------------------------------------------------------------------COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.datasets.yaml276 # ---------------------------------------------------------------------------COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.datasets.yaml278 # ---------------------------------------------------------------------------COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.datasets.yaml282 # ---------------------------------------------------------------------------COMMENT
MEDIUMtools/spicepodschema/tests/spicepod.datasets.yaml284 # ---------------------------------------------------------------------------COMMENT
491 more matches not shown…
Structural Annotation Overuse164 hits · 268 pts
SeverityFileLineSnippetContext
LOWtools/testoperator/src/commands/streaming/mutations.rs383 // Step 1: INSERT wrong values for all mutated rowsCOMMENT
LOWtools/testoperator/src/commands/streaming/mutations.rs405 // Step 2: UPDATE correct values for update-path rowsCOMMENT
LOWtools/testoperator/src/commands/streaming/mutations.rs420 // Step 3: DELETE delete-path rowsCOMMENT
LOWtools/testoperator/src/commands/streaming/mutations.rs432 // Step 4: INSERT correct values for delete-path rowsCOMMENT
LOWtools/testoperator/src/commands/streaming/mutations.rs447 // Step 5: Direct INSERT remaining rowsCOMMENT
LOW…ator-cayenne/src/partitioned_insert_strategy/insert.rs123 // Step 1: route each input batch to its partition's writer task.COMMENT
LOW…ator-cayenne/src/partitioned_insert_strategy/insert.rs180 // Step 2: join every writer task, collecting prepared overwrites andCOMMENT
LOW…ator-cayenne/src/partitioned_insert_strategy/insert.rs204 // Step 3: catalog transaction. Open once, apply every partition'sCOMMENT
LOW…ator-cayenne/src/partitioned_insert_strategy/insert.rs222 // Step 4: per-partition in-memory finish (snapshot id, listingCOMMENT
LOW…es/runtime-acceleration/benches/snapshot_operations.rs516 // Step 1: Copy file locallyCOMMENT
LOW…es/runtime-acceleration/benches/snapshot_operations.rs522 // Step 2: Read and compute checksumCOMMENT
LOW…es/runtime-acceleration/benches/snapshot_operations.rs526 // Step 3: Upload to S3COMMENT
LOW…es/runtime-acceleration/benches/snapshot_operations.rs534 // Step 4: Update metadata in S3COMMENT
LOW…es/runtime-acceleration/benches/snapshot_operations.rs689 // Step 1: Copy file locallyCOMMENT
LOW…es/runtime-acceleration/benches/snapshot_operations.rs695 // Step 2: Read and compute checksumCOMMENT
LOW…es/runtime-acceleration/benches/snapshot_operations.rs699 // Step 3: Upload to S3COMMENT
LOW…es/runtime-acceleration/benches/snapshot_operations.rs707 // Step 4: Update metadata in S3COMMENT
LOWcrates/runtime-acceleration/src/snapshot/mod.rs1269 // Step 0: Engine-specific live checkpoint while the lock is held.COMMENT
LOWcrates/runtime-acceleration/src/snapshot/mod.rs1278 // Step 1: Copy the database file locally (lock is held)COMMENT
LOWcrates/runtime-acceleration/src/snapshot/mod.rs1287 // Step 2: Release the lock - queries can resumeCOMMENT
LOWcrates/runtime-acceleration/src/snapshot/mod.rs1294 // Step 3: Prepare a snapshot using engine-specific logicCOMMENT
LOWcrates/runtime-acceleration/src/snapshot/mod.rs1301 // Step 4: Upload the file (with retry for transient network errors)COMMENT
LOWcrates/runtime-acceleration/src/snapshot/mod.rs1376 // Step 2: Release the lock - queries can resumeCOMMENT
LOWcrates/runtime-acceleration/src/snapshot/mod.rs1383 // Step 3: Upload the tar archive (with retry for transient network errors)COMMENT
LOWcrates/runtime-acceleration/src/snapshot/mod.rs1321 // Step 5: Cleanup temp filesCOMMENT
LOWcrates/runtime-acceleration/src/snapshot/mod.rs1341 // Step 0: Ask the engine for any per-directory skip list / extras.COMMENT
LOWcrates/runtime-acceleration/src/snapshot/mod.rs1354 // Step 1: Create a temporary tar archive of all directoriesCOMMENT
LOWcrates/runtime-acceleration/src/snapshot/mod.rs1403 // Step 4: Cleanup temp archiveCOMMENT
LOW…es/cayenne/tests/upsert_with_pending_deletions_test.rs418 // Step 1: Initial insert - PKs 0, 1, 2COMMENT
LOW…es/cayenne/tests/upsert_with_pending_deletions_test.rs439 // Step 2: Upsert PK 2 + insert NEW PK 3COMMENT
LOW…es/cayenne/tests/upsert_with_pending_deletions_test.rs461 // Step 3: Insert PK 3 again + NEW PK 4COMMENT
LOW…es/cayenne/tests/upsert_with_pending_deletions_test.rs546 // Step 1: Initial insert - PKs 1, 2, 3COMMENT
LOW…es/cayenne/tests/upsert_with_pending_deletions_test.rs567 // Step 2: Upsert PK 2 - this creates pending deletionsCOMMENT
LOW…es/cayenne/tests/upsert_with_pending_deletions_test.rs586 // Step 3: Insert NEW PK 4 while pending deletions existCOMMENT
LOW…es/cayenne/tests/upsert_with_pending_deletions_test.rs601 // Step 4: Query all rows - protected snapshot data must be visibleCOMMENT
LOW…es/cayenne/tests/upsert_with_pending_deletions_test.rs664 // Step 1: Initial insert - PKs 1, 2, 3COMMENT
LOW…es/cayenne/tests/upsert_with_pending_deletions_test.rs686 // Step 2: Upsert key=2 - creates pending deletion + protected snapshot S_ACOMMENT
LOW…es/cayenne/tests/upsert_with_pending_deletions_test.rs707 // Step 3: Explicit DELETE key=2 - THE CRUCIAL STEPCOMMENT
LOW…es/cayenne/tests/upsert_with_pending_deletions_test.rs722 // Step 4: Upsert key=2 again - BEFORE THE FIX, this would NOT detect theCOMMENT
LOW…es/cayenne/tests/upsert_with_pending_deletions_test.rs745 // Step 5: Upsert key=2 one more time - before the fix, this created yetCOMMENT
LOWcrates/cayenne/tests/staged_append_test.rs490 // Step 1: Insert initial data so snapshot directory existsCOMMENT
LOWcrates/cayenne/tests/staged_append_test.rs497 // Step 2: Corrupt the target snapshot directory — replace it with a regular file.COMMENT
LOWcrates/cayenne/tests/staged_append_test.rs509 // Step 3: Attempt another insert — should fail during the move phase.COMMENT
LOWcrates/cayenne/tests/staged_append_test.rs533 // Step 4: Verify WAL persists in an isolated staging dir after the failed move.COMMENT
LOWcrates/cayenne/tests/data_inlining_test.rs1252 // Step 1: 8 batches above INLINE_MAX_ROWS so each writes a Vortex fileCOMMENT
LOW…es/cayenne/tests/protected_snapshot_projection_test.rs111 // Step 1: Insert initial data (ids 1, 2, 3)COMMENT
LOW…es/cayenne/tests/protected_snapshot_projection_test.rs122 // Step 2: Delete id=2 to create pending deletions (deletion seq=1)COMMENT
LOW…es/cayenne/tests/protected_snapshot_projection_test.rs126 // Step 3: Insert new data (ids 4, 5) - creates protected snapshot with max_delete_seq=1COMMENT
LOW…es/cayenne/tests/protected_snapshot_projection_test.rs137 // Step 4: Delete id=4 AFTER the protected snapshot was created (deletion seq=2)COMMENT
LOW…es/cayenne/tests/protected_snapshot_projection_test.rs144 // Step 5: Query with projection that EXCLUDES the PK column (id)COMMENT
LOWcrates/cayenne/tests/large_upsert_test.rs86 // Step 1: Insert ROW_COUNT rows (ids 0..ROW_COUNT).COMMENT
LOWcrates/cayenne/tests/large_upsert_test.rs98 // Step 2: Upsert the same ROW_COUNT PKs with updated values.COMMENT
LOWcrates/cayenne/tests/large_upsert_test.rs110 // Step 3: Verify all rows have updated values.COMMENT
LOWcrates/cayenne/tests/large_upsert_test.rs187 // Step 1: Insert ROW_COUNT distinct PKs.COMMENT
LOWcrates/cayenne/tests/large_upsert_test.rs199 // Step 2: Upsert the same PKs with new values (exercises the keyedCOMMENT
LOWcrates/cayenne/tests/large_upsert_test.rs211 // Step 3: Exactly ROW_COUNT live rows — no PK dropped or duplicated by theCOMMENT
LOWcrates/cayenne/tests/large_upsert_test.rs224 // Step 4: Every PK maps to its updated value (correct routing + latest-wins).COMMENT
LOWcrates/cayenne/src/ddl/physical_plans.rs564 // Step 1: Execute the join plan to get matched rows with updated values.COMMENT
LOWcrates/cayenne/src/ddl/physical_plans.rs587 // Step 2: Validate no duplicate target keys in join output.COMMENT
LOWcrates/cayenne/src/ddl/physical_plans.rs623 // Step 5: Insert updated rows into the target.COMMENT
104 more matches not shown…
Verbosity Indicators148 hits · 244 pts
SeverityFileLineSnippetContext
LOWtools/testoperator/src/commands/streaming/mutations.rs383 // Step 1: INSERT wrong values for all mutated rowsCOMMENT
LOWtools/testoperator/src/commands/streaming/mutations.rs405 // Step 2: UPDATE correct values for update-path rowsCOMMENT
LOWtools/testoperator/src/commands/streaming/mutations.rs420 // Step 3: DELETE delete-path rowsCOMMENT
LOWtools/testoperator/src/commands/streaming/mutations.rs432 // Step 4: INSERT correct values for delete-path rowsCOMMENT
LOWtools/testoperator/src/commands/streaming/mutations.rs447 // Step 5: Direct INSERT remaining rowsCOMMENT
LOWcrates/accelerators/accelerator-cayenne/src/lib.rs3265 // For S3 Express, we need to check if the metadata database exists locallyCOMMENT
LOW…ator-cayenne/src/partitioned_insert_strategy/insert.rs123 // Step 1: route each input batch to its partition's writer task.COMMENT
LOW…ator-cayenne/src/partitioned_insert_strategy/insert.rs180 // Step 2: join every writer task, collecting prepared overwrites andCOMMENT
LOW…ator-cayenne/src/partitioned_insert_strategy/insert.rs204 // Step 3: catalog transaction. Open once, apply every partition'sCOMMENT
LOW…ator-cayenne/src/partitioned_insert_strategy/insert.rs222 // Step 4: per-partition in-memory finish (snapshot id, listingCOMMENT
LOW…es/runtime-acceleration/benches/snapshot_operations.rs516 // Step 1: Copy file locallyCOMMENT
LOW…es/runtime-acceleration/benches/snapshot_operations.rs522 // Step 2: Read and compute checksumCOMMENT
LOW…es/runtime-acceleration/benches/snapshot_operations.rs526 // Step 3: Upload to S3COMMENT
LOW…es/runtime-acceleration/benches/snapshot_operations.rs534 // Step 4: Update metadata in S3COMMENT
LOW…es/runtime-acceleration/benches/snapshot_operations.rs689 // Step 1: Copy file locallyCOMMENT
LOW…es/runtime-acceleration/benches/snapshot_operations.rs695 // Step 2: Read and compute checksumCOMMENT
LOW…es/runtime-acceleration/benches/snapshot_operations.rs699 // Step 3: Upload to S3COMMENT
LOW…es/runtime-acceleration/benches/snapshot_operations.rs707 // Step 4: Update metadata in S3COMMENT
LOWcrates/runtime-acceleration/src/snapshot/mod.rs1269 // Step 0: Engine-specific live checkpoint while the lock is held.COMMENT
LOWcrates/runtime-acceleration/src/snapshot/mod.rs1278 // Step 1: Copy the database file locally (lock is held)COMMENT
LOWcrates/runtime-acceleration/src/snapshot/mod.rs1287 // Step 2: Release the lock - queries can resumeCOMMENT
LOWcrates/runtime-acceleration/src/snapshot/mod.rs1294 // Step 3: Prepare a snapshot using engine-specific logicCOMMENT
LOWcrates/runtime-acceleration/src/snapshot/mod.rs1301 // Step 4: Upload the file (with retry for transient network errors)COMMENT
LOWcrates/runtime-acceleration/src/snapshot/mod.rs1376 // Step 2: Release the lock - queries can resumeCOMMENT
LOWcrates/runtime-acceleration/src/snapshot/mod.rs1383 // Step 3: Upload the tar archive (with retry for transient network errors)COMMENT
LOWcrates/runtime-acceleration/src/snapshot/mod.rs1321 // Step 5: Cleanup temp filesCOMMENT
LOWcrates/runtime-acceleration/src/snapshot/mod.rs1341 // Step 0: Ask the engine for any per-directory skip list / extras.COMMENT
LOWcrates/runtime-acceleration/src/snapshot/mod.rs1354 // Step 1: Create a temporary tar archive of all directoriesCOMMENT
LOWcrates/runtime-acceleration/src/snapshot/mod.rs1403 // Step 4: Cleanup temp archiveCOMMENT
LOW…es/cayenne/tests/upsert_with_pending_deletions_test.rs418 // Step 1: Initial insert - PKs 0, 1, 2COMMENT
LOW…es/cayenne/tests/upsert_with_pending_deletions_test.rs439 // Step 2: Upsert PK 2 + insert NEW PK 3COMMENT
LOW…es/cayenne/tests/upsert_with_pending_deletions_test.rs461 // Step 3: Insert PK 3 again + NEW PK 4COMMENT
LOW…es/cayenne/tests/upsert_with_pending_deletions_test.rs546 // Step 1: Initial insert - PKs 1, 2, 3COMMENT
LOW…es/cayenne/tests/upsert_with_pending_deletions_test.rs567 // Step 2: Upsert PK 2 - this creates pending deletionsCOMMENT
LOW…es/cayenne/tests/upsert_with_pending_deletions_test.rs586 // Step 3: Insert NEW PK 4 while pending deletions existCOMMENT
LOW…es/cayenne/tests/upsert_with_pending_deletions_test.rs601 // Step 4: Query all rows - protected snapshot data must be visibleCOMMENT
LOW…es/cayenne/tests/upsert_with_pending_deletions_test.rs664 // Step 1: Initial insert - PKs 1, 2, 3COMMENT
LOW…es/cayenne/tests/upsert_with_pending_deletions_test.rs686 // Step 2: Upsert key=2 - creates pending deletion + protected snapshot S_ACOMMENT
LOW…es/cayenne/tests/upsert_with_pending_deletions_test.rs707 // Step 3: Explicit DELETE key=2 - THE CRUCIAL STEPCOMMENT
LOW…es/cayenne/tests/upsert_with_pending_deletions_test.rs722 // Step 4: Upsert key=2 again - BEFORE THE FIX, this would NOT detect theCOMMENT
LOW…es/cayenne/tests/upsert_with_pending_deletions_test.rs745 // Step 5: Upsert key=2 one more time - before the fix, this created yetCOMMENT
LOWcrates/cayenne/tests/staged_append_test.rs490 // Step 1: Insert initial data so snapshot directory existsCOMMENT
LOWcrates/cayenne/tests/staged_append_test.rs497 // Step 2: Corrupt the target snapshot directory — replace it with a regular file.COMMENT
LOWcrates/cayenne/tests/staged_append_test.rs509 // Step 3: Attempt another insert — should fail during the move phase.COMMENT
LOWcrates/cayenne/tests/staged_append_test.rs533 // Step 4: Verify WAL persists in an isolated staging dir after the failed move.COMMENT
LOWcrates/cayenne/tests/data_inlining_test.rs1252 // Step 1: 8 batches above INLINE_MAX_ROWS so each writes a Vortex fileCOMMENT
LOW…es/cayenne/tests/protected_snapshot_projection_test.rs111 // Step 1: Insert initial data (ids 1, 2, 3)COMMENT
LOW…es/cayenne/tests/protected_snapshot_projection_test.rs122 // Step 2: Delete id=2 to create pending deletions (deletion seq=1)COMMENT
LOW…es/cayenne/tests/protected_snapshot_projection_test.rs126 // Step 3: Insert new data (ids 4, 5) - creates protected snapshot with max_delete_seq=1COMMENT
LOW…es/cayenne/tests/protected_snapshot_projection_test.rs137 // Step 4: Delete id=4 AFTER the protected snapshot was created (deletion seq=2)COMMENT
LOW…es/cayenne/tests/protected_snapshot_projection_test.rs144 // Step 5: Query with projection that EXCLUDES the PK column (id)COMMENT
LOWcrates/cayenne/tests/large_upsert_test.rs86 // Step 1: Insert ROW_COUNT rows (ids 0..ROW_COUNT).COMMENT
LOWcrates/cayenne/tests/large_upsert_test.rs98 // Step 2: Upsert the same ROW_COUNT PKs with updated values.COMMENT
LOWcrates/cayenne/tests/large_upsert_test.rs110 // Step 3: Verify all rows have updated values.COMMENT
LOWcrates/cayenne/tests/large_upsert_test.rs187 // Step 1: Insert ROW_COUNT distinct PKs.COMMENT
LOWcrates/cayenne/tests/large_upsert_test.rs199 // Step 2: Upsert the same PKs with new values (exercises the keyedCOMMENT
LOWcrates/cayenne/tests/large_upsert_test.rs211 // Step 3: Exactly ROW_COUNT live rows — no PK dropped or duplicated by theCOMMENT
LOWcrates/cayenne/tests/large_upsert_test.rs224 // Step 4: Every PK maps to its updated value (correct routing + latest-wins).COMMENT
LOWcrates/cayenne/src/ddl/physical_plans.rs564 // Step 1: Execute the join plan to get matched rows with updated values.COMMENT
LOWcrates/cayenne/src/ddl/physical_plans.rs587 // Step 2: Validate no duplicate target keys in join output.COMMENT
88 more matches not shown…
Hyper-Verbose Identifiers153 hits · 153 pts
SeverityFileLineSnippetContext
LOW…picepods/chbench/local-writeback/direct_write_bench.py243def check_writeback_converges(base_url, pg_dsn, w_id, timeout_s):CODE
LOWtest/tpc-bench/scylladb/setup_tpch.py21def create_keyspace_and_tables(session):CODE
LOWtest/adbc/python/test_flightsql_adbc.py91def test_prepared_statement_simple(conn) -> bool:CODE
LOWtest/adbc/python/test_flightsql_adbc.py121def test_prepared_statement_multiple_params(conn) -> bool:CODE
LOWtest/adbc/python/test_flightsql_adbc.py153def test_prepared_statement_with_strings(conn) -> bool:CODE
LOWtest/adbc/python/test_flightsql_adbc.py183def test_prepared_statement_types(conn) -> bool:CODE
LOWtest/adbc/python/test_flightsql_adbc.py218def test_prepared_statement_null(conn) -> bool:CODE
LOWtest/adbc/python/test_flightsql_adbc.py251def test_prepared_statement_reuse(conn) -> bool:CODE
LOWtest/adbc/python/test_flightsql_adbc.py288def test_prepare_execute_commit_pattern(conn) -> bool:CODE
LOWexamples/runtime_demo.py111def simulate_runtime_arrow_mem(user_id):STRING
LOWexamples/runtime_demo.py138def simulate_concurrent_queries(num_users, function_name):STRING
LOWscripts/load_tpch_scylladb.sh332def convert_value_for_resource(value, dtype):CODE
LOWscripts/check_rust_gate_paths.py118def extract_code_change_globs() -> list[str]:CODE
LOWscripts/check_sccache_call_sites.py29def forwards_to_setup_sccache(action_yml):CODE
LOW.github/scripts/test_check_workflow_yaml.py139 def test_a_step_budget_below_the_jobs_is_accepted(self):CODE
LOW.github/scripts/test_check_workflow_yaml.py144 def test_a_step_budget_equal_to_the_jobs_is_reported(self):CODE
LOW.github/scripts/test_check_workflow_yaml.py153 def test_a_step_budget_above_the_jobs_is_reported(self):CODE
LOW.github/scripts/test_check_workflow_yaml.py162 def test_a_step_budget_without_a_job_budget_is_left_alone(self):CODE
LOW.github/scripts/test_check_workflow_yaml.py168 def test_a_job_budget_without_step_budgets_is_left_alone(self):CODE
LOW.github/scripts/test_check_workflow_yaml.py173 def test_an_unnamed_step_is_reported_by_position(self):CODE
LOW.github/scripts/test_check_workflow_yaml.py214 def test_the_first_suite_needs_no_condition(self):CODE
LOW.github/scripts/test_check_workflow_yaml.py220 def test_a_later_suite_with_no_condition_is_reported(self):CODE
LOW.github/scripts/test_check_workflow_yaml.py228 def test_a_later_suite_gated_only_on_a_secret_is_reported(self):CODE
LOW.github/scripts/test_check_workflow_yaml.py237 def test_every_hidden_suite_is_reported(self):CODE
LOW.github/scripts/test_check_workflow_yaml.py245 def test_a_not_cancelled_guard_is_accepted(self):CODE
LOW.github/scripts/test_check_workflow_yaml.py288 def test_a_pr_gate_is_left_alone(self):CODE
LOW.github/scripts/test_check_workflow_yaml.py297 def test_a_step_that_runs_no_suite_is_left_alone(self):CODE
LOW.github/scripts/test_check_workflow_yaml.py307 def test_a_snapshot_push_behind_a_suite_is_reported(self):CODE
LOW.github/scripts/test_check_workflow_yaml.py347 def test_secrets_in_a_step_condition_is_reported_with_the_remedy(self):CODE
LOW.github/scripts/test_check_workflow_yaml.py356 def test_secrets_in_a_job_condition_is_reported(self):CODE
LOW.github/scripts/test_check_workflow_yaml.py363 def test_bracket_property_access_is_reported_like_dotted_access(self):CODE
LOW.github/scripts/test_check_workflow_yaml.py371 def test_bracket_access_in_a_step_condition_is_reported_with_the_remedy(self):CODE
LOW.github/scripts/test_check_workflow_yaml.py379 def test_bracket_access_on_an_available_context_is_accepted(self):CODE
LOW.github/scripts/test_check_workflow_yaml.py389 def test_a_context_name_inside_a_quoted_literal_is_not_a_reference(self):CODE
LOW.github/scripts/test_check_workflow_yaml.py404 def test_a_real_reference_beside_a_quoted_literal_is_still_reported(self):CODE
LOW.github/scripts/test_check_workflow_yaml.py414 def test_the_env_indirection_trunk_uses_is_accepted(self):CODE
LOW.github/scripts/test_check_workflow_yaml.py422 def test_env_is_rejected_in_a_job_condition_but_allowed_in_a_step_condition(self):CODE
LOW.github/scripts/test_check_workflow_yaml.py475 def test_a_dotted_path_after_an_allowed_context_is_not_read_as_a_context(self):CODE
LOW.github/scripts/test_check_workflow_yaml.py485 def test_every_unavailable_context_in_one_condition_is_named_once(self):CODE
LOW.github/scripts/test_check_workflow_yaml.py492 def test_an_unknown_dotted_word_is_left_alone(self):CODE
LOW.github/scripts/test_check_workflow_yaml.py502 def test_a_condition_without_any_if_key_is_not_invented(self):CODE
LOW.github/scripts/test_check_workflow_yaml.py505 def test_secrets_in_a_composite_action_step_condition_is_reported(self):CODE
LOW.github/scripts/test_check_workflow_yaml.py558 def test_a_test_run_after_the_cache_setup_is_reported(self):CODE
LOW.github/scripts/test_check_workflow_yaml.py565 def test_dropping_the_endpoint_first_is_accepted(self):CODE
LOW.github/scripts/test_check_workflow_yaml.py573 def test_a_spice_run_is_covered_too(self):CODE
LOW.github/scripts/test_check_workflow_yaml.py582 def test_a_job_without_the_cache_setup_is_left_alone(self):CODE
LOW.github/scripts/test_check_workflow_yaml.py590 def test_building_the_archive_still_gets_the_cache(self):CODE
LOW.github/scripts/test_check_workflow_yaml.py684 def test_valid_action_has_no_problems(self):CODE
LOW.github/scripts/test_check_workflow_yaml.py687 def test_missing_runs_block_is_reported(self):CODE
LOW.github/scripts/test_check_workflow_yaml.py691 def test_invalid_yaml_is_reported(self):CODE
LOW.github/scripts/test_check_workflow_yaml.py696 def test_an_action_is_not_required_to_declare_triggers_or_jobs(self):CODE
LOW.github/scripts/test_check_workflow_yaml.py256 def test_an_always_guard_is_accepted(self):CODE
LOW.github/scripts/test_check_workflow_yaml.py265 def test_a_status_function_inside_a_quoted_literal_does_not_count(self):CODE
LOW.github/scripts/test_check_workflow_yaml.py276 def test_a_real_status_function_beside_a_quoted_literal_is_accepted(self):CODE
LOW.github/scripts/test_check_workflow_yaml.py322 def test_an_unnamed_suite_is_reported_by_position(self):CODE
LOW.github/scripts/test_check_workflow_yaml.py332def _workflow_with_conditions(job_if: str = "", step_if: str = "") -> str:CODE
LOW.github/scripts/test_check_workflow_yaml.py436 def test_steps_matrix_and_runner_are_rejected_in_a_job_condition(self):CODE
LOW.github/scripts/test_check_workflow_yaml.py449 def test_contexts_every_condition_may_name_are_accepted(self):CODE
LOW.github/scripts/test_check_workflow_yaml.py464 def test_a_hyphenated_job_name_is_not_read_as_a_context(self):CODE
LOW.github/scripts/test_check_workflow_yaml.py516 def test_a_composite_action_step_may_name_inputs_steps_and_runner(self):CODE
93 more matches not shown…
AI Slop Vocabulary43 hits · 126 pts
SeverityFileLineSnippetContext
MEDIUM.config/nextest.toml7# Tests must be written to be thread-safe and robust under maximum parallelism.COMMENT
MEDIUMcrates/runtime-tls/src/reload.rs616 // Watching the parent dir is more robust toCOMMENT
MEDIUM…rs/connector-sharepoint/src/sharepoint/object_store.rs623/// Simpler and more robust than trying to stream SharePoint's upload-sessionCOMMENT
MEDIUMcrates/cache/src/lib.rs171/// until TTL expiry. Resolving both sides first makes the comparison robust toCOMMENT
MEDIUMcrates/accelerators/accelerator-cayenne/src/imds.rs116/// but the digit guard keeps it robust to future families).COMMENT
MEDIUMcrates/cayenne/tests/delta_encoding_test.rs51/// so the level-0 vs full-level on-disk size gap is large and robust.COMMENT
MEDIUMcrates/cayenne/tests/retention_test.rs672 // A 1-hour window keeps the test robust to wall-clock execution time (theCOMMENT
MEDIUMcrates/cayenne/benches/vs_duckdb_memory_budget.rs48//! The high-water lines are order-robust; the wall-clock lanes are theCOMMENT
MEDIUMcrates/cayenne/src/lib.rs51//! This design allows Cayenne to leverage Vortex's columnar format and `DataFusion`'sCOMMENT
MEDIUMcrates/cayenne/src/provider/tuning.rs1185 /// carried a source-commit ts. The freshness-goal control/SLO signal — robust toCOMMENT
MEDIUMcrates/cayenne/src/provider/tuning.rs2459 // are tried in a robust, observable order rather than from per-write phaseCOMMENT
MEDIUMcrates/cayenne/src/provider/context.rs887 // share the robust signal. Falls back to the instantaneous age until theCOMMENT
MEDIUMcrates/runtime/tests/acceleration/localpod_sync.rs356 // robust against that subscription landing slightly after readiness, without dependingCOMMENT
MEDIUMcrates/runtime/tests/snapshot_refresh/mod.rs418/// avoid `SELECT count(*)` so the assertion is robust across engines thatCOMMENT
LOWcrates/runtime/tests/kafka/test_data/stack_qa.json15{"question_id":73841495,"title":"Why is this C while didn't working anymore?","question_html":"<pre><code>#include&lt;stCODE
MEDIUMcrates/runtime/src/secrets_preflight.rs319 // JSON so the test stays robust to field additions.COMMENT
MEDIUMcrates/runtime-datafusion-udfs/src/digest_many.rs232 // Hash entire array in one call - hash function can leverage SIMD internallyCOMMENT
MEDIUMcrates/data_components/src/turso.rs179/// While the current implementation is robust, consider using **RFC3339 TEXT format** if:COMMENT
MEDIUMcrates/dataformat-json/src/stream.rs68/// * Uses `serde_json` for robust parsing.COMMENT
MEDIUMcrates/llms/tests/llms/mod.rs681 // JSON Parse the function arguments to ensure robust to ordering.COMMENT
MEDIUMcrates/document_parse/src/pdfium.rs26//! access. This is the robust path for locked-down / air-gapped containers.COMMENT
MEDIUMcrates/spice-table/src/lib.rs170 /// This is deliberately robust to `filters` referencing columns this index knows nothingCOMMENT
MEDIUM…earch/financebench/full_text_search-cayenne[file].yaml45 # under. `content AS text` matches the column name the search harness reports on. AccelerationCOMMENT
MEDIUMtest/spicepods/search/custom/example.yaml6# `--benchmark-dataset`. The harness tests the three fixed tables it queries — `corpus`,COMMENT
MEDIUMtest/spicepods/search/custom/example.yaml8# fields at your own sources; `views:` map real column names onto the names/types the harness needs.COMMENT
MEDIUMtest/spicepods/search/custom/example.yaml30 # Query set: the harness reads `SELECT _id as id, text FROM test_queries`.COMMENT
MEDIUMtest/spicepods/search/custom/example.yaml34 # Relevance judgments (qrels): the harness readsCOMMENT
MEDIUM…picepods/chbench/local-writeback/direct_write_bench.py196 # / per-table-OCC-starvation scenario the harness is meant to catch.COMMENT
MEDIUM…h/accelerated/postgres-cayenne[file]-full-refresh.yaml6# (no OLTP load) for robust explain-plan snapshots for analytical queries.COMMENT
MEDIUMdocs/cayenne/package-lock.json3008 "resolved": "https://registry.npmjs.org/robust-predicates/-/robust-predicates-3.0.3.tgz",CODE
MEDIUMscripts/ci-search-test-report.py320 # (the harness always prints this after all the per-test "---- name stdout ----"COMMENT
MEDIUMscripts/test_preflight_build_disk.sh167# about the subject finding no `df`, not about the harness finding no shell.COMMENT
MEDIUMscripts/htap-explain-probe.sh110# Probe until the associated spiced exits. The parent harness normally kills thisCOMMENT
MEDIUMscripts/test_check_rust_gate_paths.py225 # never read, so this harness must not pass on the same condition either.COMMENT
MEDIUMscripts/loaded_mutation_property_test.sh9# isolated run can pass indefinitely). This harness recreates that pressure byCOMMENT
MEDIUMscripts/test_restack_stacked_branch.sh36# a counter incremented in there is discarded on the way out -- a harness thatCOMMENT
MEDIUMscripts/verify_cli_build.py49 # harness. Both are reported with name `spice` and kind `["bin"]`, andCOMMENT
MEDIUMscripts/verify_cli_build.py50 # only `profile.test` tells them apart — the harness answers `--version`COMMENT
MEDIUM.github/workflows/testoperator_run_bench.yml502 # default with a one-second `item_ttl`, and the harness runs one warmup queryCOMMENT
MEDIUM.github/scripts/chbench_template.sh11# Flow (mirrors the local harness `do_reset`):COMMENT
MEDIUM.github/actions/setup-chbench-mysql/action.yml18 # ("Cannot initialize AIO sub-system"). Simulated I/O is robust and theCOMMENT
MEDIUM.github/actions/check-code-changes/action.yml68 # The reachability guard's own regression harness runs in `lint-rust`,COMMENT
MEDIUM.github/prompts/scripts/check_ste.py21# commit. `--json` prints machine-readable metrics for the eval harness.COMMENT
Modern AI Meta-Vocabulary16 hits · 48 pts
SeverityFileLineSnippetContext
MEDIUMREADME.md184### Retrieval-Augmented Generation (RAG)COMMENT
MEDIUMtools/testoperator/README.md211##### Run a HTTP consistency test against an embedding modelCOMMENT
MEDIUMtools/testoperator/README.md257##### Run a HTTP overhead test against an embedding modelCOMMENT
MEDIUMtools/spicepodschema/README.md85 ├── main.rs # Entry point, orchestrationCODE
MEDIUMtools/spicepodschema/tests/spicepod.embeddings.yaml3# EMBEDDINGS - Test coverage for embedding model sourcesCOMMENT
MEDIUMtools/spidapter/README.md341##### Run a HTTP consistency test against an embedding modelCOMMENT
MEDIUMtools/spidapter/README.md387##### Run a HTTP overhead test against an embedding modelCOMMENT
MEDIUMcrates/runtime/src/lib.rs212 #[snafu(display("Failed to initialize embedding model: {source}"))]COMMENT
MEDIUMcrates/runtime-search/src/rerank.rs411/// The UDTF scaffold. Analogous to [`VectorSearchTableFunc`] — holds a weakCOMMENT
MEDIUMcrates/embed-api/src/lib.rs78 #[snafu(display("Failed to load embedding model: {source}"))]COMMENT
MEDIUMcrates/llms/src/chat/mistral.rs577 // mistral.rs v0.9.0 agentic / code-execution / shell / files features:COMMENT
MEDIUMcrates/llms/src/chat/mistral.rs865 // mistral.rs v0.9.0 agentic / diffusion / file responses: Spice doesCOMMENT
MEDIUMbin/spice/src/commands/search.rs43 spice search --model my_embed # Use a specific embedding modelCODE
MEDIUM.github/workflows/e2e_test_ci.yml396 # chat model and the embedding model with one mechanism, and an unset secretCOMMENT
MEDIUM.github/workflows/e2e_test_ci.yml409 # chat model and the embedding model are on disk, later runs need no HubCOMMENT
MEDIUM.github/workflows/testoperator_dispatch_search.yml80 # that share an embedding model provider throttle each other when run concurrentlyCOMMENT
Self-Referential Comments13 hits · 38 pts
SeverityFileLineSnippetContext
MEDIUM…/runtime/tests/tpcds_postgres/setup_local_test_data.sh221 # Create the bucket (ignore error if it already exists)COMMENT
MEDIUM…s/data_components/tests/hadoop_data/docker-compose.yml1# Create a MinIO cluster, so we can copy the hadoop warehouse into it for testingCOMMENT
MEDIUMinstall/install.sh145 # Create the temp directoryCOMMENT
MEDIUMinstall/install-spiced.sh217 # Create the temp directoryCOMMENT
MEDIUM…ch/sf1/accelerated/constraints/simulate_append_data.sh13# Create the directory if it does not existCOMMENT
MEDIUMtest/scripts/install-scripts-test.sh100# Create a temporary directory for test artifactsCOMMENT
MEDIUMtest/scripts/clickbench/setup-file.bash5# Create the folderCOMMENT
MEDIUMscripts/test_reclaim_runner_disk.sh86 '# This file is a cache directory tag created by cargo.' >"$dir/CACHEDIR.TAG"CODE
MEDIUMscripts/test_reclaim_runner_disk.sh159 '# This file is a cache directory tag automatically created by restic.' \CODE
MEDIUMscripts/load_tpch_scylladb.sh189 # Create a Python script to load data via boto3COMMENT
MEDIUMscripts/load_tpch_scylladb.sh434 # Create a virtual environment if boto3 is not availableCOMMENT
MEDIUMscripts/reclaim_runner_disk.sh101readonly CARGO_PRODUCER_LINE='# This file is a cache directory tag created by cargo.'CODE
MEDIUM.github/workflows/build_and_release_cuda.yml282 # Create an array of actual files (no globs)COMMENT
Redundant / Tautological Comments25 hits · 37 pts
SeverityFileLineSnippetContext
LOWtools/spidapter/scenarios/mongodb-streams.yaml17# Set MONGO_SPICEPOD_PATH to an ABSOLUTE path to a pods/mongo/*.yaml — in the spidapterCOMMENT
LOW…/runtime/tests/tpcds_postgres/setup_local_test_data.sh117 # Check if container already existsCOMMENT
LOWinstall/install-build.sh173# Check if a string looks like a commit SHA (7+ hex characters)COMMENT
LOWinstall/install-build.sh216 # Check if this run has the artifact we needCOMMENT
LOWinstall/install-build.sh273 # Check if this run has the artifact we needCOMMENT
LOWinstall/install.sh415 # Check if running interactivelyCOMMENT
LOWinstall/install.sh451 # Check if PATH is already configured properly (look for actual export/setenv/fish_add_path commands)COMMENT
LOWinstall/install-nightly.sh229 # Check if the response contains an error message (run not found, etc.)COMMENT
LOWinstall/install-nightly.sh238 # Check if artifacts array existsCOMMENT
LOWinstall/install-nightly.sh443 # Check if PATH already contains the install directoryCOMMENT
LOW…/tpch/sf1/federated/dynamodb[scylladb-alternator].yaml21 # Set aws_auth to 'key' with dummy values or configure Alternator authCOMMENT
LOWtest/scripts/install-scripts-test.sh679 # Check if a known artifact URL returns 302 (redirect to download)COMMENT
LOWtest/adbc/python/run_test.sh67# Check if uv or python3 is availableCOMMENT
LOWtest/adbc/python/run_test.sh91# Check if dependencies are installed (only when not using uv run)COMMENT
LOWtest/adbc/python/run_test.sh110 # Check if spiced is in PATH or use the one in ~/.spice/binCOMMENT
LOWscripts/distributed.sh34# Check if spice CLI is availableCOMMENT
LOWscripts/distributed.sh40# Check if spiced binary existsCOMMENT
LOWscripts/distributed.sh65 # Check if certificate already existsCOMMENT
LOWscripts/load_tpch_scylladb.sh183 # Check if Python and boto3 are availableCOMMENT
LOWscripts/preflight_build_disk.sh25# Set it to 0 to report the readings and refuse nothing,COMMENT
LOW.github/workflows/e2e_test_release_install_helm.yml302 # Check if release is deployedCOMMENT
LOW.github/workflows/build_nightly.yml51 # Check if there are new commits since the last nightly tagCOMMENT
LOW.github/scripts/validate_table_providers_commit.sh56# Check if the commit exists and is reachable from the spiceai branchCOMMENT
LOW.github/scripts/get_release_version.py14 # Set LATEST_RELEASE to trueCOMMENT
LOW.github/actions/setup-make/action.yml33 # Check if make is already installedCOMMENT
Fake / Example Data29 hits · 30 pts
SeverityFileLineSnippetContext
LOWcrates/data-connectors/connector-dynamodb/src/schema.rs818 ("name".to_string(), av_string("John Doe")),CODE
LOWcrates/data-connectors/connector-dynamodb/src/schema.rs1141 contact_map.insert("phone".to_string(), av_string("555-1234"));CODE
LOW…ates/runtime/tests/cluster/distributed_acceleration.rs211 (1,'John Doe',28,'New York',85),(2,'Jane Smith',34,'Los Angeles',92),CODE
LOWcrates/runtime/tests/hashicorp_vault/mod.rs100 (1, 'Acme Corp'), (2, 'Globex'), (3, 'Initech');",CODE
LOWcrates/runtime/tests/hashicorp_vault/mod.rs142 assert_eq!(customers.value(0), "Acme Corp");CODE
LOWcrates/runtime/tests/http/json_nested_fields.rs59 "actor": "admin@example.com",CODE
LOWcrates/runtime/tests/http/json_nested_fields.rs89 "actor": "admin@example.com",CODE
LOWcrates/runtime/tests/http/json_nested_fields.rs90 "approvedBy": "admin@example.com",CODE
LOWcrates/runtime/tests/kafka/test_data/orders_nested.json8 "contact": { "email": "alice@example.com", "phone": "555-1234" }CODE
LOWcrates/runtime/tests/kafka/test_data/orders_nested.json53 "contact": { "email": "dana@example.org" }CODE
LOWcrates/runtime/tests/graphql/mod.rs79 name: "John Doe".to_string(),CODE
LOWcrates/runtime/tests/graphql/mod.rs95 name: "Jane Doe".to_string(),CODE
LOWcrates/runtime/src/dataconnector/sink.rs50 "placeholder",CODE
LOWcrates/runtime/src/dataconnector/deferred.rs44 "placeholder",CODE
LOWcrates/chunking/benches/chunking.rs45 Lorem ipsum dolor sit amet, consectetur adipiscing elit. Sed do eiusmod tempor incididunt ut \CODE
LOWcrates/chunking/benches/chunking.rs45 Lorem ipsum dolor sit amet, consectetur adipiscing elit. Sed do eiusmod tempor incididunt ut \CODE
LOWcrates/data_components/src/object/mod.rs190 (r"^.*([a-z]+)@.*$", "user@example.com"), // Character class with quantifierCODE
LOWcrates/dataformat-json/src/stream.rs1541 assert_eq!(filter_to_string(input), r#"{"name": "John Doe"}"#);CODE
LOWcrates/dataformat-json/src/unnest.rs1532 "phone": "555-1234"CODE
LOWcrates/dataformat-json/src/unnest.rs1549 "user.details.profile.contact.phone": "555-1234",CODE
LOWcrates/arrow_tools/src/format.rs1040 "Lorem ipsum dolor sit amet, consectetur adipiscing elit.CODE
LOWcrates/arrow_tools/src/format.rs1040 "Lorem ipsum dolor sit amet, consectetur adipiscing elit.CODE
LOWcrates/arrow_tools/src/format.rs1047 "Lorem ipsum dolor adipiscing elit.CODE
LOWtest/payloads/model-generic-lorem.txt1Lorem ipsum dolor sit amet, consectetur adipiscing elit. Curabitur eget tortor vitae nulla tincidunt fermentum.CODE
LOWtest/payloads/model-generic-lorem.txt1Lorem ipsum dolor sit amet, consectetur adipiscing elit. Curabitur eget tortor vitae nulla tincidunt fermentum.CODE
LOWdocs/release_notes/v1.6/v1.6.1.md44 "phone": "555-1234"CODE
LOWdocs/threat_models/v2.0.0.json16 "placeholder": "New STRIDE diagram description",CODE
LOWdocs/threat_models/v1.9.1.json16 "placeholder": "New STRIDE diagram description",CODE
LOWdocs/threat_models/v0.17.4-beta.json16 "placeholder": "New STRIDE diagram description",CODE
Example Usage Blocks19 hits · 28 pts
SeverityFileLineSnippetContext
LOW…/runtime/tests/tpcds_postgres/setup_local_test_data.sh27# Usage:COMMENT
LOWinstall/install-build.sh5# Usage:COMMENT
LOWtest/scripts/install-scripts-test.sh11# Usage:COMMENT
LOWtest/adbc/python/run_test.sh4# Usage:COMMENT
LOWdeploy/chart/examples/aws-irsa-values.yaml14# Usage:COMMENT
LOWscripts/tpcds_explain.sh6# Usage:COMMENT
LOWscripts/check_crate_layers.py18# Usage:COMMENT
LOWscripts/ci-search-test-report.py18# Usage:COMMENT
LOWscripts/htap-explain-probe.sh19# Usage:COMMENT
LOWscripts/restack_stacked_branch.sh33# Usage:COMMENT
LOWscripts/check_fork_patches.py23# Usage:COMMENT
LOWscripts/check_rust_gate_paths.py30# Usage:COMMENT
LOWscripts/run-chbench-htap.sh14# Usage:COMMENT
LOWscripts/loaded_mutation_property_test.sh17# Usage:COMMENT
LOWscripts/chbench-waterfall.py41# Usage:COMMENT
LOWscripts/check_module_reachability.py19# Usage:COMMENT
LOW.github/scripts/reap_superseded_merge_queue_runs.py26# Usage:COMMENT
LOW.github/prompts/evals/score_eval.py17# Usage:COMMENT
LOW.github/prompts/scripts/check_ste.py23# Usage:COMMENT
Excessive Try-Catch Wrapping18 hits · 19 pts
SeverityFileLineSnippetContext
LOWtest/adbc/python/README.md207 except Exception as e:CODE
MEDIUMtest/adbc/python/test_flightsql_adbc.py28 print(f"Error: Missing required package: {e}")CODE
LOWtest/adbc/python/test_flightsql_adbc.py86 except Exception as e:CODE
LOWtest/adbc/python/test_flightsql_adbc.py116 except Exception as e:CODE
LOWtest/adbc/python/test_flightsql_adbc.py148 except Exception as e:CODE
LOWtest/adbc/python/test_flightsql_adbc.py178 except Exception as e:CODE
LOWtest/adbc/python/test_flightsql_adbc.py213 except Exception as e:CODE
LOWtest/adbc/python/test_flightsql_adbc.py246 except Exception as e:CODE
LOWtest/adbc/python/test_flightsql_adbc.py283 except Exception as e:CODE
LOWtest/adbc/python/test_flightsql_adbc.py340 except Exception as e:CODE
LOWtest/adbc/python/test_flightsql_adbc.py348 except Exception as e:CODE
LOWtest/adbc/python/test_flightsql_adbc.py391 except Exception as e:CODE
LOWtest/adbc/python/test_flightsql_adbc.py410 except Exception as e:CODE
LOWdocs/cayenne/gen_waterfall.py121except Exception as e: # noqa: BLE001 — preview is best-effortCODE
LOWscripts/load_tpch_scylladb.sh327 except Exception as e:CODE
LOWscripts/load_tpch_scylladb.sh375 except Exception as e:CODE
LOWscripts/load_tpch_scylladb.sh390 except Exception:CODE
LOWscripts/load_tpch_scylladb.sh404 except Exception as e:CODE
Unused Imports14 hits · 14 pts
SeverityFileLineSnippetContext
LOWscripts/check_crate_layers.py25CODE
LOWscripts/ci-search-test-report.py28CODE
LOWscripts/check_table_layers.py28CODE
LOWscripts/test_check_module_reachability.py13CODE
LOWscripts/test_check_rust_gate_paths.py15CODE
LOWscripts/check_fork_patches.py29CODE
LOWscripts/financebench_stage.py49CODE
LOWscripts/check_rust_gate_paths.py35CODE
LOWscripts/test_check_fork_patches.py19CODE
LOWscripts/check_module_reachability.py25CODE
LOW.github/scripts/check_nextest_config.py33CODE
LOW.github/scripts/check_workflow_yaml.py36CODE
LOW.github/prompts/evals/score_eval.py21CODE
LOW.github/prompts/scripts/check_ste.py27CODE
Deep Nesting13 hits · 13 pts
SeverityFileLineSnippetContext
LOWcrates/runtime/tests/file/generate_parquet.py4CODE
LOWscripts/check_crate_layers.py111CODE
LOWscripts/ci-search-test-report.py199CODE
LOWscripts/ci-search-test-report.py253CODE
LOWscripts/ci-search-test-report.py296CODE
LOWscripts/ci-search-test-report.py454CODE
LOWscripts/generate_changelog.py18CODE
LOWscripts/financebench_stage.py259CODE
LOWscripts/chbench-waterfall.py461CODE
LOWscripts/check_module_reachability.py92CODE
LOWscripts/check_module_reachability.py188CODE
LOWscripts/check_module_reachability.py284CODE
LOW.github/prompts/scripts/check_ste.py207CODE
AI Response Leakage2 hits · 12 pts
SeverityFileLineSnippetContext
HIGHcrates/runtime/tests/kafka/test_data/stack_qa.json30{"question_id":73841031,"title":"How to test the error message from anyhow::Error?","question_html":"<p>There is <code>cCODE
HIGHcrates/runtime/tests/kafka/test_data/stack_qa.json80{"question_id":73839994,"title":"Website Layout height problems what am i doing wrong?","question_html":"<p><strong>HellCODE
Magic Placeholder Names2 hits · 10 pts
SeverityFileLineSnippetContext
HIGHdocs/release_notes/v1.8/v1.8.0.md185spice chat --cloud --api-key <your-api-key>CODE
HIGHdocs/release_notes/v1.8/v1.8.0.md186spice search --cloud --api-key <your-api-key>CODE
Slop Phrases4 hits · 10 pts
SeverityFileLineSnippetContext
MEDIUMcrates/runtime/tests/kafka/test_data/stack_qa.json66{"question_id":73840195,"title":"Symfony like view helpers in laravel?","question_html":"<p>Symfony has this:</p>\n<p><aCODE
LOWcrates/runtime/tests/kafka/test_data/stack_qa.json72{"question_id":73840129,"title":"how to use only strings to pull json data","question_html":"<p>So I keep running into aCODE
MEDIUMcrates/runtime/tests/kafka/test_data/stack_qa.json89{"question_id":73839832,"title":"TypeScript: How to refine a type where one of it's properties is probably null","questiCODE
MEDIUMcrates/llms/src/rerank/mod.rs401 "Sure! Here is the ranking:\n[{\"id\":1,\"score\":0.7}]\nLet me know if you need more.";CODE