Repository Analysis

OpenBMB/MiniCPM

MiniCPM5-1B: A SOTA 1B on-device LLM, small yet powerful.

3.9 Likely human-written View on GitHub

Analysis Overview

This report presents the forensic synthetic code analysis of OpenBMB/MiniCPM, a Jupyter Notebook project with 10,087 GitHub stars. SynthScan v2.0 examined 45,489 lines of code across 150 source files, recording 126 pattern matches distributed across 10 syntactic categories. The overall adjusted score of 3.9 places this repository in the Likely human-written band.

The scanner applied 160+ deterministic lexical heuristics, multi-line block detectors, abstract syntax tree depth profilers, and a cross-file Jaccard similarity matrix to construct a statistically normalised synthetic code estimate. All matches are individually weighted by severity coefficient and contextual multiplier before summation, and the resulting headline score is temporally discounted to account for the repository's development history relative to the commercial emergence of large language model coding tooling (November 2022 onward).

3.9
Adjusted Score
3.9
Raw Score
100%
Time Factor
2026-07-27
Last Push
10.1K
Stars
Jupyter Notebook
Language
45.5K
Lines of Code
150
Files
126
Pattern Hits
2026-08-02
Scan Date
0.00
HC Hit Rate

What These Metrics Mean

Adjusted Score
Primary synthetic code indicator. Raw score normalised per 1,000 lines of code and multiplied by the temporal discount factor. This is the definitive comparative metric — use it to rank repositories by AI authorship density.
Raw Score
The unmodified sum of all severity-weighted, context-multiplied pattern match scores before temporal discounting. Reflects the absolute signal strength independent of when the repository was last active.
Time Factor
The temporal discount multiplier (0–100%) applied to the raw score. Repositories last updated before ChatGPT's launch (Nov 2022) receive a 5% factor. Full signal is only assigned to repositories active in the post-adoption era (Jan 2024+).
Pattern Hits
Total count of individual pattern matches across all files and categories. A high hit count with a low score may indicate a very large codebase with isolated AI snippets; a low count with a high score indicates dense, concentrated AI signatures.
HC Hit Rate
High+Critical pattern hits per file, averaged across the repository. This orthogonal signal catches repositories where a few files are densely packed with high-severity AI tells — a strong indicator even when the normalised score appears moderate due to codebase size.
Lines of Code / Files
Total lines and files analysed. The scanner examines 94 file extensions. These denominators are used to normalise the score, enabling fair comparison between repositories of vastly different sizes.

Score History

Longitudinal tracking requires multiple scan runs. Once this repository is re-scanned after new commits land, this chart will visualise how the synthetic code signal evolves over time — enabling you to detect whether AI authorship is growing, stabilising, or being actively corrected by human engineers.

No multi-scan history yet — run the scanner again to build trend data.

Severity Breakdown

Classifies detected patterns by their diagnostic confidence and structural impact. CRITICAL patterns (coefficient 10) represent definitive synthetic signatures — hallucinated imports, explicit LLM attribution metadata — virtually never produced by human authors. HIGH (5) indicates strong structural tells such as cross-file repetition or cross-linguistic idioms. MEDIUM (2) covers recognisable conversational padding and AI-specific vocabulary. LOW (1) captures subtle indicators like tautological comments and generic boilerplate that require density to carry independent signal.

CRITICAL 0HIGH 0MEDIUM 45LOW 81

Directory Score Breakdown

This horizontal bar chart decomposes the repository's raw synthetic code score by top-level directory, allowing you to pinpoint precisely which modules or components carry the highest AI authorship density. Directories with disproportionately high scores relative to their size warrant targeted manual review: concentrated AI signatures often trace back to mass-generated configuration layers, auto-ported test suites, LLM-scaffolded boilerplate classes, or entire subsystems authored under heavy copilot assistance. Use this view to prioritise your human code-review effort.

Pattern Findings

The scanner identified 126 distinct pattern matches across 10 syntactic categories. Each entry below represents a discrete location in the source code where the engine recorded a statistically significant AI authorship indicator. Expand any category row to inspect the individual file paths, line numbers, code snippets, and the lexical context (CODE, COMMENT, or STRING) in which each match was detected.

Reading the findings table: The Severity column indicates the diagnostic confidence level (CRITICAL / HIGH / MEDIUM / LOW). The Context column identifies whether the match occurred inside executable code, an inline comment, or a string literal — comment-context matches receive a ×1.5 weight because LLMs systematically over-annotate. The ⚡ bolt icon marks clustered matches: three or more patterns within a 10-line window, each receiving an additional ×1.5 density multiplier as dense clusters constitute far stronger evidence of synthetic authorship than isolated hits.

Slop Phrases32 hits · 61 pts
SeverityFileLineSnippetContext
MEDIUM…ma_factory_example/llama_factory_data/kto_en_demo.json67 "content": "Certainly, measuring the results of a marketing strategy and making adjustments based on those metriCODE
MEDIUM…ma_factory_example/llama_factory_data/kto_en_demo.json394 "content": "Using Unreal's Gameplay Ability System (GAS) in conjunction with interpolation and prediction can prCODE
LOW…ma_factory_example/llama_factory_data/kto_en_demo.json903 "content": "I apologize for the confusion. The `unique_id` parameter in the `momepy.Tessellation` function expecCODE
MEDIUM…ma_factory_example/llama_factory_data/kto_en_demo.json2012 "content": "Two-factor authentication (2FA) adds an extra layer of security by requiring users to provide two foCODE
MEDIUM…ma_factory_example/llama_factory_data/kto_en_demo.json2114 "content": "To adapt the program to handle in-place modification or the creation of new objects based on a booleCODE
MEDIUM…ma_factory_example/llama_factory_data/kto_en_demo.json2449 "content": "I'm not sure where you've heard that information, but as an AI language model, I can share some factCODE
MEDIUM…ma_factory_example/llama_factory_data/kto_en_demo.json2790 "content": "To derive the original matrix from the rotated one, you would need to perform a 90 degrees counter-cCODE
MEDIUM…ma_factory_example/llama_factory_data/kto_en_demo.json3013 "content": "The genetic factors that contribute to the intricate design of the peacock's tail feathers are quiteCODE
MEDIUM…ma_factory_example/llama_factory_data/kto_en_demo.json3585 "content": "As an AI language model,I cannot determine the sentiment of the product review without further conteCODE
MEDIUM…ma_factory_example/llama_factory_data/kto_en_demo.json3611 "content": "I'm unable to provide you with a real-time percentage of population living in poverty in MarlboroughCODE
LOW…ma_factory_example/llama_factory_data/kto_en_demo.json3690 "content": "Firstly, as a fitness trainer, I highly recommend consulting with a healthcare professional before eCODE
MEDIUM…ma_factory_example/llama_factory_data/kto_en_demo.json3850 "content": "Since each medium-sized box measures 2 feet on all sides, its volume will be: 2*2*2 = 8 cubic feet. CODE
MEDIUM…ma_factory_example/llama_factory_data/kto_en_demo.json4774 "content": "In Feng Shui, the five elements—wood, fire, earth, metal, and water—are believed to interact in certCODE
MEDIUM…ma_factory_example/llama_factory_data/kto_en_demo.json4897 "content": "Sure, I understand the problem. However, I'm a large language model, I cannot provide a correct answCODE
MEDIUM…ma_factory_example/llama_factory_data/kto_en_demo.json4923 "content": "As an AI language model, I am not able to provide program codes or algorithms. However, I can providCODE
MEDIUM…ma_factory_example/llama_factory_data/sft_zh_demo.json4470 "output": "As an AI language model I am capable of analyzing text and identifying emotions, however, the sentence prCODE
MEDIUM…tune/llama_factory/llama_factory_data/kto_en_demo.json67 "content": "Certainly, measuring the results of a marketing strategy and making adjustments based on those metriCODE
MEDIUM…tune/llama_factory/llama_factory_data/kto_en_demo.json394 "content": "Using Unreal's Gameplay Ability System (GAS) in conjunction with interpolation and prediction can prCODE
LOW…tune/llama_factory/llama_factory_data/kto_en_demo.json903 "content": "I apologize for the confusion. The `unique_id` parameter in the `momepy.Tessellation` function expecCODE
MEDIUM…tune/llama_factory/llama_factory_data/kto_en_demo.json2012 "content": "Two-factor authentication (2FA) adds an extra layer of security by requiring users to provide two foCODE
MEDIUM…tune/llama_factory/llama_factory_data/kto_en_demo.json2114 "content": "To adapt the program to handle in-place modification or the creation of new objects based on a booleCODE
MEDIUM…tune/llama_factory/llama_factory_data/kto_en_demo.json2449 "content": "I'm not sure where you've heard that information, but as an AI language model, I can share some factCODE
MEDIUM…tune/llama_factory/llama_factory_data/kto_en_demo.json2790 "content": "To derive the original matrix from the rotated one, you would need to perform a 90 degrees counter-cCODE
MEDIUM…tune/llama_factory/llama_factory_data/kto_en_demo.json3013 "content": "The genetic factors that contribute to the intricate design of the peacock's tail feathers are quiteCODE
MEDIUM…tune/llama_factory/llama_factory_data/kto_en_demo.json3585 "content": "As an AI language model,I cannot determine the sentiment of the product review without further conteCODE
MEDIUM…tune/llama_factory/llama_factory_data/kto_en_demo.json3611 "content": "I'm unable to provide you with a real-time percentage of population living in poverty in MarlboroughCODE
LOW…tune/llama_factory/llama_factory_data/kto_en_demo.json3690 "content": "Firstly, as a fitness trainer, I highly recommend consulting with a healthcare professional before eCODE
MEDIUM…tune/llama_factory/llama_factory_data/kto_en_demo.json3850 "content": "Since each medium-sized box measures 2 feet on all sides, its volume will be: 2*2*2 = 8 cubic feet. CODE
MEDIUM…tune/llama_factory/llama_factory_data/kto_en_demo.json4774 "content": "In Feng Shui, the five elements—wood, fire, earth, metal, and water—are believed to interact in certCODE
MEDIUM…tune/llama_factory/llama_factory_data/kto_en_demo.json4897 "content": "Sure, I understand the problem. However, I'm a large language model, I cannot provide a correct answCODE
MEDIUM…tune/llama_factory/llama_factory_data/kto_en_demo.json4923 "content": "As an AI language model, I am not able to provide program codes or algorithms. However, I can providCODE
MEDIUM…tune/llama_factory/llama_factory_data/sft_zh_demo.json4470 "output": "As an AI language model I am capable of analyzing text and identifying emotions, however, the sentence prCODE
Self-Referential Comments10 hits · 26 pts
SeverityFileLineSnippetContext
MEDIUMdemo/minicpm3/code_interpreter/code_interpreter.py105 # Create a SamplingParams objectSTRING
MEDIUMdemo/minicpm3/code_interpreter/code_interpreter.py135 # Define a regular expression pattern to match Python code blocksCOMMENT
MEDIUM…ma_factory_example/llama_factory_data/kto_en_demo.json887 "content": "To smoothen the output polygons without creating gaps or overlaps, you can use a combination of librCODE
MEDIUM…ma_factory_example/llama_factory_data/kto_en_demo.json895 "content": "I apologize for the confusion. Let's try a different approach using the `TopologicalPreserveSimplifiCODE
MEDIUM…ma_factory_example/llama_factory_data/kto_en_demo.json907 "content": "I have the following code:\n#%% Generate geoTiffs with index values based on one large geotiff with CODE
MEDIUM…ma_factory_example/llama_factory_data/kto_en_demo.json919 "content": "To smoothen the output polygons without creating gaps or overlays between the polygons, you can use CODE
MEDIUM…tune/llama_factory/llama_factory_data/kto_en_demo.json887 "content": "To smoothen the output polygons without creating gaps or overlaps, you can use a combination of librCODE
MEDIUM…tune/llama_factory/llama_factory_data/kto_en_demo.json895 "content": "I apologize for the confusion. Let's try a different approach using the `TopologicalPreserveSimplifiCODE
MEDIUM…tune/llama_factory/llama_factory_data/kto_en_demo.json907 "content": "I have the following code:\n#%% Generate geoTiffs with index values based on one large geotiff with CODE
MEDIUM…tune/llama_factory/llama_factory_data/kto_en_demo.json919 "content": "To smoothen the output polygons without creating gaps or overlays between the polygons, you can use CODE
Excessive Try-Catch Wrapping21 hits · 26 pts
SeverityFileLineSnippetContext
LOWdemo/minicpm4/MCP/generate_example.py63 except Exception as e:CODE
LOWdemo/minicpm4/MCP/eval_scripts.py83 except Exception as e:CODE
LOW…o/minicpm4/SurveyGeneration/src/retriever/retriever.py163 except Exception as e:CODE
MEDIUM…o/minicpm4/SurveyGeneration/src/retriever/retriever.py164 print(f"Error in call_search_engine: {e}")CODE
MEDIUM…o/minicpm4/SurveyGeneration/src/retriever/retriever.py133def call_search_engine(tool_call, topk=10):CODE
LOWdemo/minicpm4/SurveyGeneration/src/generation/run.py51 except Exception as e:CODE
MEDIUMdemo/minicpm4/SurveyGeneration/src/generation/run.py52 print(f"Error sending to WebSocket: {e}")CODE
LOWdemo/minicpm4/SurveyGeneration/src/generation/run.py210 except Exception as e:CODE
MEDIUMdemo/minicpm4/SurveyGeneration/src/generation/run.py211 print(f"Error posting to frontend: {e}")CODE
LOWdemo/minicpm4/SurveyGeneration/src/generation/run.py310 except Exception as e:CODE
MEDIUMdemo/minicpm4/SurveyGeneration/src/generation/run.py311 print(f"Error generating survey: {e}")CODE
LOWdemo/minicpm3/function_call/minicpm_tool_parser.py205 except Exception as e:CODE
LOWdemo/minicpm3/function_call/minicpm_tool_parser.py277 except Exception as e:CODE
LOWdemo/minicpm3/code_interpreter/code_interpreter.py76 except Exception as e:STRING
LOWtool_parsers/minicpm5xml_tool_parser.py40except Exception: # pragma: no coverCODE
LOWtool_parsers/minicpm5xml_tool_parser.py108 except Exception:CODE
LOWtool_parsers/minicpm5xml_tool_parser.py226 except Exception:CODE
LOWtool_parsers/minicpm5xml_tool_parser.py260 except Exception:CODE
LOWtool_parsers/minicpm5xml_tool_parser.py402 except Exception as e:CODE
LOWtool_parsers/minicpm5xml_tool_parser.py645 except Exception:CODE
LOWfinetune/mlx_finetune.py441 except Exception as e:CODE
Deep Nesting25 hits · 25 pts
SeverityFileLineSnippetContext
LOWdemo/minicpm4/MCP/generate_example.py25CODE
LOWdemo/minicpm4/MCP/generate_example.py71CODE
LOWdemo/minicpm4/MCP/eval_scripts.py21CODE
LOWdemo/minicpm4/SurveyGeneration/src/generation/run.py93CODE
LOWdemo/minicpm4/SurveyGeneration/src/generation/run.py227CODE
LOWdemo/minicpm4/SurveyGeneration/src/generation/run.py255CODE
LOWdemo/minicpm4/SurveyGeneration/src/generation/buffer.py36CODE
LOWdemo/minicpm4/SurveyGeneration/src/generation/buffer.py75CODE
LOWdemo/minicpm4/SurveyGeneration/src/generation/buffer.py149CODE
LOWdemo/minicpm4/SurveyGeneration/src/generation/buffer.py210CODE
LOWdemo/minicpm4/SurveyGeneration/src/generation/buffer.py408CODE
LOWdemo/minicpm4/SurveyGeneration/src/generation/buffer.py462CODE
LOWdemo/minicpm4/SurveyGeneration/src/generation/buffer.py510CODE
LOWdemo/minicpm4/SurveyGeneration/src/generation/buffer.py556CODE
LOWdemo/minicpm4/SurveyGeneration/src/generation/buffer.py612CODE
LOWdemo/minicpm4/SurveyGeneration/src/generation/buffer.py639CODE
LOWdemo/minicpm3/function_call/minicpm_tool_parser.py150CODE
LOWdemo/minicpm3/function_call/minicpm_tool_parser.py239CODE
LOWdemo/minicpm3/function_call/minicpm_tool_parser.py75CODE
LOWtool_parsers/minicpm5xml_tool_parser.py161CODE
LOWfinetune/finetune.py78CODE
LOWfinetune/sft_dpo_trainer/finetune_dpo_trainer.py85CODE
LOWfinetune/sft_dpo_trainer/finetune_dpo_trainer.py184CODE
LOWfinetune/sft_dpo_trainer/finetune_dpo_trainer.py228CODE
LOWminicpm_sala/finetune/trainer/finetune.py78CODE
Unused Imports17 hits · 17 pts
SeverityFileLineSnippetContext
LOWdemo/minicpm/langchain_demo.py35CODE
LOWdemo/minicpm/langchain_demo.py40CODE
LOWdemo/minicpm4/MCP/eval_scripts.py1CODE
LOW…inicpm4/SurveyGeneration/src/preprocess/build_index.py6CODE
LOW…inicpm4/SurveyGeneration/src/preprocess/build_index.py8CODE
LOWdemo/minicpm4/SurveyGeneration/src/generation/buffer.py276CODE
LOWdemo/minicpm3/code_interpreter/code_interpreter.py3CODE
LOWdemo/minicpm3/code_interpreter/code_interpreter.py4CODE
LOWdemo/minicpm3/code_interpreter/code_interpreter.py6CODE
LOWdemo/minicpm3/code_interpreter/code_interpreter.py7CODE
LOWtool_parsers/minicpm5xml_tool_parser.py6CODE
LOWquantize/awq_quantize.py4CODE
LOWquantize/quantize_eval.py5CODE
LOWfinetune/mlx_finetune.py46CODE
LOWfinetune/sft_dpo_trainer/finetune_dpo_trainer.py5CODE
LOWfinetune/sft_dpo_trainer/finetune_dpo_trainer.py5CODE
LOWfinetune/sft_dpo_trainer/finetune_dpo_trainer.py19CODE
Hyper-Verbose Identifiers10 hits · 10 pts
SeverityFileLineSnippetContext
LOWdemo/minicpm4/SurveyGeneration/src/generation/buffer.py75 def convert_survey_dict_to_str(current_survey):CODE
LOWdemo/minicpm4/SurveyGeneration/src/generation/buffer.py149 def convert_survey_dict_to_abbr_str(current_survey):CODE
LOWdemo/minicpm4/SurveyGeneration/src/generation/buffer.py408 def _build_user_prompt_force_correct(query, current_survey, trajs):CODE
LOWdemo/minicpm4/SurveyGeneration/src/generation/buffer.py510 def build_prompt_for_generator(self):CODE
LOWdemo/minicpm3/function_call/minicpm_tool_parser.py75 def extract_tool_calls_streaming(CODE
LOWtool_parsers/minicpm5xml_tool_parser.py464 def _process_complete_block_streaming(CODE
LOWtool_parsers/minicpm5xml_tool_parser.py526 def _process_partial_block_streaming(CODE
LOWtool_parsers/minicpm5xml_tool_parser.py565 def extract_tool_calls_streaming(CODE
LOWquantize/quantize_data/alpaca_data_cleaned.json730 "output": "Here's a Python implementation of the function:\n\n```\ndef uppercase_unless_rejected(string):\n if stCODE
LOWfinetune/sft_dpo_trainer/finetune_dpo_trainer.py228 def encode_conversation_with_labels(self, messages, response):CODE
Redundant / Tautological Comments4 hits · 5 pts
SeverityFileLineSnippetContext
LOWdemo/minicpm3/code_interpreter/code_interpreter.py165 # Check if the response contains the termination keywordCOMMENT
LOWfinetune/mlx_finetune.py487 # Check if any sequence is longer than 2048 tokensCOMMENT
LOW…ma_factory_example/llama_factory_data/sft_zh_demo.json275 "output": "def calculate_sum(numbers):\n if not isinstance(numbers, list): # Check if input is a list\n rCODE
LOW…tune/llama_factory/llama_factory_data/sft_zh_demo.json275 "output": "def calculate_sum(numbers):\n if not isinstance(numbers, list): # Check if input is a list\n rCODE
Over-Commented Block4 hits · 4 pts
SeverityFileLineSnippetContext
LOW…/SurveyGeneration/frontend/minicpm4-survey/src/App.jsx181 };COMMENT
LOWdemo/minicpm4/SurveyGeneration/src/generation/buffer.py1COMMENT
LOWdemo/minicpm3/function_call/function_calling.py41 # "tool_calls": [COMMENT
LOWdemo/minicpm3/function_call/function_calling.py61 # {COMMENT
AI Slop Vocabulary2 hits · 4 pts
SeverityFileLineSnippetContext
MEDIUM…ma_factory_example/llama_factory_data/kto_en_demo.json2769 "content": " As an Azure Cloud Engineer working with Microsoft Azure, you can indeed utilize Azure Active DirectCODE
MEDIUM…tune/llama_factory/llama_factory_data/kto_en_demo.json2769 "content": " As an Azure Cloud Engineer working with Microsoft Azure, you can indeed utilize Azure Active DirectCODE
AI Structural Patterns1 hit · 1 pts
SeverityFileLineSnippetContext
LOW…o/minicpm4/SurveyGeneration/src/retriever/retriever.py131CODE