Speech recognition module for Python, supporting several engines and APIs, online and offline.
This report presents the forensic synthetic code analysis of Uberi/speech_recognition, a Python project with 8,987 GitHub stars. SynthScan v2.0 examined 5,971 lines of code across 61 source files, recording 76 pattern matches distributed across 6 syntactic categories. The overall adjusted score of 14.0 places this repository in the Low AI signal band.
The scanner applied 160+ deterministic lexical heuristics, multi-line block detectors, abstract syntax tree depth profilers, and a cross-file Jaccard similarity matrix to construct a statistically normalised synthetic code estimate. All matches are individually weighted by severity coefficient and contextual multiplier before summation, and the resulting headline score is temporally discounted to account for the repository's development history relative to the commercial emergence of large language model coding tooling (November 2022 onward).
Longitudinal tracking requires multiple scan runs. Once this repository is re-scanned after new commits land, this chart will visualise how the synthetic code signal evolves over time — enabling you to detect whether AI authorship is growing, stabilising, or being actively corrected by human engineers.
Classifies detected patterns by their diagnostic confidence and structural impact. CRITICAL patterns (coefficient 10) represent definitive synthetic signatures — hallucinated imports, explicit LLM attribution metadata — virtually never produced by human authors. HIGH (5) indicates strong structural tells such as cross-file repetition or cross-linguistic idioms. MEDIUM (2) covers recognisable conversational padding and AI-specific vocabulary. LOW (1) captures subtle indicators like tautological comments and generic boilerplate that require density to carry independent signal.
This horizontal bar chart decomposes the repository's raw synthetic code score by top-level directory, allowing you to pinpoint precisely which modules or components carry the highest AI authorship density. Directories with disproportionately high scores relative to their size warrant targeted manual review: concentrated AI signatures often trace back to mass-generated configuration layers, auto-ported test suites, LLM-scaffolded boilerplate classes, or entire subsystems authored under heavy copilot assistance. Use this view to prioritise your human code-review effort.
The scanner identified 76 distinct pattern matches across 6 syntactic categories. Each entry below represents a discrete location in the source code where the engine recorded a statistically significant AI authorship indicator. Expand any category row to inspect the individual file paths, line numbers, code snippets, and the lexical context (CODE, COMMENT, or STRING) in which each match was detected.
Reading the findings table: The Severity column indicates the diagnostic confidence level (CRITICAL / HIGH / MEDIUM / LOW). The Context column identifies whether the match occurred inside executable code, an inline comment, or a string literal — comment-context matches receive a ×1.5 weight because LLMs systematically over-annotate. The ⚡ bolt icon marks clustered matches: three or more patterns within a 10-line window, each receiving an additional ×1.5 density multiplier as dense clusters constitute far stronger evidence of synthetic authorship than isolated hits.
| Severity | File | Line | Snippet | Context |
|---|---|---|---|---|
| LOW | tests/test_recognition.py | 20 | def test_recognizer_attributes(self): | CODE |
| LOW⚡ | tests/test_audio.py | 133 | def test_returns_self_when_already_fits(self): | CODE |
| LOW⚡ | tests/test_audio.py | 139 | def test_raises_when_max_bytes_too_small(self): | CODE |
| LOW⚡ | tests/test_audio.py | 144 | def test_raises_on_unaligned_frame_data(self): | CODE |
| LOW⚡ | tests/test_audio.py | 154 | def test_raises_when_max_bytes_below_one_sample_per_width(self): | CODE |
| LOW | tests/test_audio.py | 174 | def test_fixed_split_chunks_fit_within_max_bytes(self): | CODE |
| LOW | tests/test_audio.py | 191 | def test_fixed_split_aligns_to_sample_boundary(self): | CODE |
| LOW | tests/test_audio.py | 198 | def test_silence_aware_raises_setup_error_without_librosa(self): | CODE |
| LOW | tests/test_audio.py | 216 | def test_silence_aware_translates_call_time_errors_to_setup_error(self): | CODE |
| LOW | tests/test_audio.py | 238 | def test_silence_aware_translates_lazy_runtime_errors_to_setup_error(self): | CODE |
| LOW | tests/test_audio.py | 282 | def test_to_float_ndarray_normalizes_each_sample_width(self): | CODE |
| LOW | tests/test_audio.py | 304 | def test_to_float_ndarray_decodes_little_endian_regardless_of_host(self): | CODE |
| LOW | tests/test_audio.py | 336 | def test_silence_aware_uses_single_nonsilent_range_boundary(self): | CODE |
| LOW | tests/test_audio.py | 368 | def test_silence_aware_respects_byte_budget_strictly(self): | CODE |
| LOW | tests/test_audio.py | 391 | def test_silence_aware_respects_byte_budget_on_realistic_audio(self): | CODE |
| LOW | tests/test_audio.py | 423 | def test_silence_aware_snaps_to_speech_end_within_lookback(self): | CODE |
| LOW | tests/test_audio.py | 464 | def test_silence_aware_splits_at_silence_boundary(self): | CODE |
| LOW | tests/recognizers/test_cohere_api.py | 12 | def test_transcribe_default_model(mock_client_cls): | CODE |
| LOW | tests/recognizers/test_cohere_api.py | 37 | def test_transcribe_with_language(mock_client_cls): | CODE |
| LOW | tests/recognizers/test_google_cloud.py | 21 | def test_transcribe_with_google_cloud_speech(SpeechClient): | CODE |
| LOW | tests/recognizers/test_google_cloud.py | 56 | def test_transcribe_with_specified_credentials(SpeechClient): | CODE |
| LOW | tests/recognizers/test_google_cloud.py | 149 | def test_transcribe_with_specified_api_parameters(SpeechClient): | CODE |
| LOW | tests/recognizers/test_vosk.py | 23 | def test_recognize_vosk_verbose(audio_data): | CODE |
| LOW | tests/recognizers/test_google.py | 68 | def test_parse_without_confidence( | CODE |
| LOW | tests/recognizers/test_google.py | 86 | def test_parse_with_confidence( | CODE |
| LOW | tests/recognizers/whisper_api/test_openai.py | 22 | def test_transcribe_with_openai_whisper(setenv_openai_api_key: None) -> None: | CODE |
| LOW | tests/recognizers/whisper_api/test_openai.py | 56 | def test_transcribe_with_gpt_transcribe(setenv_openai_api_key: None) -> None: | CODE |
| LOW | tests/recognizers/whisper_api/test_openai.py | 93 | def test_transcribe_with_specified_language(setenv_openai_api_key: None) -> None: | CODE |
| LOW | tests/recognizers/whisper_api/test_openai.py | 125 | def test_transcribe_with_specified_prompt(setenv_openai_api_key: None) -> None: | CODE |
| LOW | tests/recognizers/whisper_api/test_groq.py | 14 | def test_transcribe_with_groq_whisper(respx_mock, monkeypatch): | CODE |
| LOW | tests/recognizers/whisper_api/test_openai_compatible.py | 10 | def test_transcribe_with_openai_compatible_api(httpserver, monkeypatch): | CODE |
| LOW | speech_recognition/__init__.py | 395 | def snowboy_wait_for_hot_word(self, snowboy_location, snowboy_hot_word_files, source, timeout=None): | CODE |
| Severity | File | Line | Snippet | Context |
|---|---|---|---|---|
| LOW | tests/test_audio.py | 203 | CODE | |
| LOW | tests/test_audio.py | 223 | CODE | |
| LOW | tests/test_audio.py | 224 | CODE | |
| LOW | tests/test_audio.py | 244 | CODE | |
| LOW | tests/test_audio.py | 274 | CODE | |
| LOW | tests/test_audio.py | 275 | CODE | |
| LOW | tests/recognizers/whisper_local/test_faster_whisper.py | 1 | CODE | |
| LOW | speech_recognition/__init__.py | 5 | CODE | |
| LOW | speech_recognition/audio.py | 1 | CODE | |
| LOW | speech_recognition/recognizers/google_cloud.py | 1 | CODE | |
| LOW | speech_recognition/recognizers/pocketsphinx.py | 1 | CODE | |
| LOW | speech_recognition/recognizers/google.py | 1 | CODE | |
| LOW | speech_recognition/recognizers/cohere_api.py | 1 | CODE | |
| LOW | speech_recognition/recognizers/vosk.py | 1 | CODE | |
| LOW | speech_recognition/recognizers/whisper_local/whisper.py | 1 | CODE | |
| LOW | …ecognition/recognizers/whisper_local/faster_whisper.py | 1 | CODE | |
| LOW | speech_recognition/recognizers/whisper_local/base.py | 1 | CODE | |
| LOW | speech_recognition/recognizers/whisper_api/groq.py | 1 | CODE | |
| LOW | speech_recognition/recognizers/whisper_api/groq.py | 7 | CODE | |
| LOW | speech_recognition/recognizers/whisper_api/openai.py | 1 | CODE | |
| LOW | examples/tensorflow_commands.py | 4 | CODE |
| Severity | File | Line | Snippet | Context |
|---|---|---|---|---|
| LOW | tests/test_audio.py | 276 | except Exception as exc: | CODE |
| LOW | speech_recognition/__init__.py | 153 | except Exception: | CODE |
| LOW | speech_recognition/__init__.py | 177 | except Exception: | CODE |
| MEDIUM | speech_recognition/__init__.py | 877 | print('Error creating bucket %s: %s' % (bucket_name, exc)) | CODE |
| MEDIUM | speech_recognition/__init__.py | 896 | print('Error getting job:', exc.response) | CODE |
| LOW | speech_recognition/__init__.py | 929 | except Exception as exc: | CODE |
| LOW | speech_recognition/__init__.py | 942 | except Exception as exc: | CODE |
| MEDIUM | speech_recognition/__init__.py | 977 | print('Error starting job:', exc.response) | CODE |
| LOW | speech_recognition/audio.py | 143 | except Exception as exc: | CODE |
| LOW | speech_recognition/audio.py | 206 | except Exception as exc: | CODE |
| Severity | File | Line | Snippet | Context |
|---|---|---|---|---|
| LOW | speech_recognition/__init__.py | 128 | CODE | |
| LOW | speech_recognition/__init__.py | 231 | CODE | |
| LOW | speech_recognition/__init__.py | 466 | CODE | |
| LOW | speech_recognition/__init__.py | 572 | CODE | |
| LOW | speech_recognition/__init__.py | 825 | CODE | |
| LOW | speech_recognition/__init__.py | 585 | CODE | |
| LOW | speech_recognition/cli.py | 12 | CODE | |
| LOW | speech_recognition/audio.py | 483 | CODE | |
| LOW | speech_recognition/audio.py | 136 | CODE |
| Severity | File | Line | Snippet | Context |
|---|---|---|---|---|
| MEDIUM | speech_recognition/__init__.py | 1209 | # =============================== | COMMENT |
| MEDIUM | speech_recognition/__init__.py | 1211 | # =============================== | COMMENT |
| Severity | File | Line | Snippet | Context |
|---|---|---|---|---|
| LOW | speech_recognition/recognizers/cohere_api.py | 9 | logger = logging.getLogger(__name__) | CODE |
| LOW | speech_recognition/recognizers/whisper_api/base.py | 6 | logger = logging.getLogger(__name__) | CODE |