Repository Analysis

datawhalechina/self-llm

《开源大模型食用指南》针对中国宝宝量身打造的基于Linux环境快速微调(全参数/Lora)、部署国内外开源大模型(LLM)/多模态大模型(MLLM)教程

4.3 Likely human-written View on GitHub

Analysis Overview

This report presents the forensic synthetic code analysis of datawhalechina/self-llm, a Jupyter Notebook project with 31,276 GitHub stars. SynthScan v2.0 examined 100,364 lines of code across 375 source files, recording 257 pattern matches distributed across 14 syntactic categories. The overall adjusted score of 4.3 places this repository in the Likely human-written band.

The scanner applied 160+ deterministic lexical heuristics, multi-line block detectors, abstract syntax tree depth profilers, and a cross-file Jaccard similarity matrix to construct a statistically normalised synthetic code estimate. All matches are individually weighted by severity coefficient and contextual multiplier before summation, and the resulting headline score is temporally discounted to account for the repository's development history relative to the commercial emergence of large language model coding tooling (November 2022 onward).

4.3
Adjusted Score
4.3
Raw Score
100%
Time Factor
2026-06-17
Last Push
31.3K
Stars
Jupyter Notebook
Language
100.4K
Lines of Code
375
Files
257
Pattern Hits
2026-07-14
Scan Date
0.08
HC Hit Rate

What These Metrics Mean

Adjusted Score
Primary synthetic code indicator. Raw score normalised per 1,000 lines of code and multiplied by the temporal discount factor. This is the definitive comparative metric — use it to rank repositories by AI authorship density.
Raw Score
The unmodified sum of all severity-weighted, context-multiplied pattern match scores before temporal discounting. Reflects the absolute signal strength independent of when the repository was last active.
Time Factor
The temporal discount multiplier (0–100%) applied to the raw score. Repositories last updated before ChatGPT's launch (Nov 2022) receive a 5% factor. Full signal is only assigned to repositories active in the post-adoption era (Jan 2024+).
Pattern Hits
Total count of individual pattern matches across all files and categories. A high hit count with a low score may indicate a very large codebase with isolated AI snippets; a low count with a high score indicates dense, concentrated AI signatures.
HC Hit Rate
High+Critical pattern hits per file, averaged across the repository. This orthogonal signal catches repositories where a few files are densely packed with high-severity AI tells — a strong indicator even when the normalised score appears moderate due to codebase size.
Lines of Code / Files
Total lines and files analysed. The scanner examines 94 file extensions. These denominators are used to normalise the score, enabling fair comparison between repositories of vastly different sizes.

Score History

This chart maps the temporal evolution of the adjusted synthetic code score across successive scan runs. An upward trajectory indicates ongoing incorporation of AI-generated code or expanding LLM-assisted scaffolding; a stable or declining trajectory may reflect active human refactoring, code removal, or the adoption of stricter authorship policies. The dashed secondary line (right axis) independently tracks total raw pattern hit count, which can diverge from the normalised score when codebase size changes significantly between scans.

Severity Breakdown

Classifies detected patterns by their diagnostic confidence and structural impact. CRITICAL patterns (coefficient 10) represent definitive synthetic signatures — hallucinated imports, explicit LLM attribution metadata — virtually never produced by human authors. HIGH (5) indicates strong structural tells such as cross-file repetition or cross-linguistic idioms. MEDIUM (2) covers recognisable conversational padding and AI-specific vocabulary. LOW (1) captures subtle indicators like tautological comments and generic boilerplate that require density to carry independent signal.

CRITICAL 0HIGH 29MEDIUM 26LOW 202

Directory Score Breakdown

This horizontal bar chart decomposes the repository's raw synthetic code score by top-level directory, allowing you to pinpoint precisely which modules or components carry the highest AI authorship density. Directories with disproportionately high scores relative to their size warrant targeted manual review: concentrated AI signatures often trace back to mass-generated configuration layers, auto-ported test suites, LLM-scaffolded boilerplate classes, or entire subsystems authored under heavy copilot assistance. Use this view to prioritise your human code-review effort.

Pattern Findings

The scanner identified 257 distinct pattern matches across 14 syntactic categories. Each entry below represents a discrete location in the source code where the engine recorded a statistically significant AI authorship indicator. Expand any category row to inspect the individual file paths, line numbers, code snippets, and the lexical context (CODE, COMMENT, or STRING) in which each match was detected.

Reading the findings table: The Severity column indicates the diagnostic confidence level (CRITICAL / HIGH / MEDIUM / LOW). The Context column identifies whether the match occurred inside executable code, an inline comment, or a string literal — comment-context matches receive a ×1.5 weight because LLMs systematically over-annotate. The ⚡ bolt icon marks clustered matches: three or more patterns within a 10-line window, each receiving an additional ×1.5 density multiplier as dense clusters constitute far stronger evidence of synthetic authorship than isolated hits.

Cross-File Repetition27 hits · 135 pts
SeverityFileLineSnippetContext
HIGHmodels/ChatGLM/05-ChatGLM3-6B接入LangChain搭建知识库助手.md0使用以下上下文来回答最后的问题。如果你不知道答案,就说你不知道,不要试图编造答案。尽量使答案简明扼要。总是在回答的最后说“谢谢你的提问!”。 {context} 问题: {question} 有用的回答:STRING
HIGHmodels/InternLM/06-InternLM接入LangChain搭建知识库助手.md0使用以下上下文来回答最后的问题。如果你不知道答案,就说你不知道,不要试图编造答案。尽量使答案简明扼要。总是在回答的最后说“谢谢你的提问!”。 {context} 问题: {question} 有用的回答:STRING
HIGHmodels/Atom/03-Atom-7B-Chat 接入langchain搭建知识库助手.md0使用以下上下文来回答最后的问题。如果你不知道答案,就说你不知道,不要试图编造答案。尽量使答案简明扼要。总是在回答的最后说“谢谢你的提问!”。 {context} 问题: {question} 有用的回答:STRING
HIGHmodels/Qwen/07-Qwen-7B-Chat 接入langchain搭建知识库助手.md0使用以下上下文来回答最后的问题。如果你不知道答案,就说你不知道,不要试图编造答案。尽量使答案简明扼要。总是在回答的最后说“谢谢你的提问!”。 {context} 问题: {question} 有用的回答:STRING
HIGHmodels/Yi/02-Yi-6B-Chat 接入langchain搭建知识库助手.md0使用以下上下文来回答最后的问题。如果你不知道答案,就说你不知道,不要试图编造答案。尽量使答案简明扼要。总是在回答的最后说“谢谢你的提问!”。 {context} 问题: {question} 有用的回答:STRING
HIGHmodels/ChatGLM/05-ChatGLM3-6B接入LangChain搭建知识库助手.md0使用以下上下文来回答最后的问题。如果你不知道答案,就说你不知道,不要试图编造答 案。尽量使答案简明扼要。总是在回答的最后说“谢谢你的提问!”。 {context} 问题: {question} 有用的回答:STRING
HIGH…ChatGLM/05-ChatGLM3-6B接入LangChain搭建知识库助手/run_gradio.py0使用以下上下文来回答最后的问题。如果你不知道答案,就说你不知道,不要试图编造答 案。尽量使答案简明扼要。总是在回答的最后说“谢谢你的提问!”。 {context} 问题: {question} 有用的回答:STRING
HIGHmodels/InternLM/06-InternLM接入LangChain搭建知识库助手.md0使用以下上下文来回答最后的问题。如果你不知道答案,就说你不知道,不要试图编造答 案。尽量使答案简明扼要。总是在回答的最后说“谢谢你的提问!”。 {context} 问题: {question} 有用的回答:STRING
HIGH…s/InternLM/06-InternLM接入LangChain搭建知识库助手/run_gradio.py0使用以下上下文来回答最后的问题。如果你不知道答案,就说你不知道,不要试图编造答 案。尽量使答案简明扼要。总是在回答的最后说“谢谢你的提问!”。 {context} 问题: {question} 有用的回答:STRING
HIGHmodels/Atom/03-Atom-7B-Chat 接入langchain搭建知识库助手.md0使用以下上下文来回答最后的问题。如果你不知道答案,就说你不知道,不要试图编造答 案。尽量使答案简明扼要。总是在回答的最后说“谢谢你的提问!”。 {context} 问题: {question} 有用的回答:STRING
HIGHmodels/Qwen/07-Qwen-7B-Chat 接入langchain搭建知识库助手.md0使用以下上下文来回答最后的问题。如果你不知道答案,就说你不知道,不要试图编造答 案。尽量使答案简明扼要。总是在回答的最后说“谢谢你的提问!”。 {context} 问题: {question} 有用的回答:STRING
HIGH…/Qwen/07-Qwen-7B-Chat 接入langchain搭建知识库助手/run_gradio.py0使用以下上下文来回答最后的问题。如果你不知道答案,就说你不知道,不要试图编造答 案。尽量使答案简明扼要。总是在回答的最后说“谢谢你的提问!”。 {context} 问题: {question} 有用的回答:STRING
HIGHmodels/ChatGLM/05-ChatGLM3-6B接入LangChain搭建知识库助手.md0提醒:<br> 1. 初始化数据库时间可能较长,请耐心等待。 2. 使用中如果出现异常,将会在文本输入框进行展示,请不要惊慌。 <br>STRING
HIGH…ChatGLM/05-ChatGLM3-6B接入LangChain搭建知识库助手/run_gradio.py0提醒:<br> 1. 初始化数据库时间可能较长,请耐心等待。 2. 使用中如果出现异常,将会在文本输入框进行展示,请不要惊慌。 <br>STRING
HIGHmodels/InternLM/06-InternLM接入LangChain搭建知识库助手.md0提醒:<br> 1. 初始化数据库时间可能较长,请耐心等待。 2. 使用中如果出现异常,将会在文本输入框进行展示,请不要惊慌。 <br>STRING
HIGH…s/InternLM/06-InternLM接入LangChain搭建知识库助手/run_gradio.py0提醒:<br> 1. 初始化数据库时间可能较长,请耐心等待。 2. 使用中如果出现异常,将会在文本输入框进行展示,请不要惊慌。 <br>STRING
HIGHmodels/Atom/03-Atom-7B-Chat 接入langchain搭建知识库助手.md0提醒:<br> 1. 初始化数据库时间可能较长,请耐心等待。 2. 使用中如果出现异常,将会在文本输入框进行展示,请不要惊慌。 <br>STRING
HIGH…/Atom/03-Atom-7B-Chat 接入langchain搭建知识库助手/run_gradio.py0提醒:<br> 1. 初始化数据库时间可能较长,请耐心等待。 2. 使用中如果出现异常,将会在文本输入框进行展示,请不要惊慌。 <br>STRING
HIGHmodels/Qwen/07-Qwen-7B-Chat 接入langchain搭建知识库助手.md0提醒:<br> 1. 初始化数据库时间可能较长,请耐心等待。 2. 使用中如果出现异常,将会在文本输入框进行展示,请不要惊慌。 <br>STRING
HIGH…/Qwen/07-Qwen-7B-Chat 接入langchain搭建知识库助手/run_gradio.py0提醒:<br> 1. 初始化数据库时间可能较长,请耐心等待。 2. 使用中如果出现异常,将会在文本输入框进行展示,请不要惊慌。 <br>STRING
HIGHmodels/Gemma3/04-Gemma3-4b evalscope智商情商评测.md0以下为多个ai服务的api端点地址,用于配置任务: - siliconflow: https://api.siliconflow.cn/v1/chat/completions - dashscope: https://dashscope.aSTRING
HIGHmodels/Gemma4/04-Gemma4-E4B-it evalscope智商情商评测.md0以下为多个ai服务的api端点地址,用于配置任务: - siliconflow: https://api.siliconflow.cn/v1/chat/completions - dashscope: https://dashscope.aSTRING
HIGHmodels/Qwen3/04-Qwen3-8B EvalScope智商情商评测.md0以下为多个ai服务的api端点地址,用于配置任务: - siliconflow: https://api.siliconflow.cn/v1/chat/completions - dashscope: https://dashscope.aSTRING
HIGH…3B-Instruct/04-Hunyuan-A13B-Instruct EvalScope 并发测试.md0以下为多个ai服务的api端点地址,用于配置任务: - siliconflow: https://api.siliconflow.cn/v1/chat/completions - dashscope: https://dashscope.aSTRING
HIGHmodels/Gemma3/6-gemma3-4B-itGRPO微调及通过swanlab可视化.md0you are given a problem. think about the problem and provide your working out. place it between {reasoning_start} and {rSTRING
HIGHmodels/Gemma4/6-gemma4-E4B-itGRPO微调及通过swanlab可视化.md0you are given a problem. think about the problem and provide your working out. place it between {reasoning_start} and {rSTRING
HIGHmodels/Qwen3/10-Qwen3-8B GRPO微调及通过swanlab可视化.md0you are given a problem. think about the problem and provide your working out. place it between {reasoning_start} and {rSTRING
Excessive Try-Catch Wrapping112 hits · 122 pts
SeverityFileLineSnippetContext
MEDIUMutils.py124 print(f"Error: {data['message']}")CODE
LOWmodels_amd/gemma3/2-gemma3-4b-it 模型服务部署.md86 except Exception as e:CODE
LOWmodels_amd/qwen3/1-Qwen3-8B-AMD部署调用.md86 except Exception as e:CODE
LOWmodels_amd/qwen3/1-Qwen3-8B-AMD部署调用.md196 except Exception as e:CODE
LOWmodels_amd/qwen3/1-Qwen3-8B-AMD部署调用.md253 except Exception as e:CODE
LOWmodels_amd/qwen3/1-Qwen3-8B-AMD部署调用.md266 except Exception as e:CODE
LOWmodels_amd/qwen3/1-Qwen3-8B-AMD部署调用.md287 except Exception as e:CODE
LOWmodels_amd/qwen3/1-Qwen3-8B-AMD部署调用.md362 except Exception as e:CODE
LOWmodels/Llama4/01-Llama4-对话助手/01-Llama4-对话助手.md277 except Exception as e:CODE
LOWmodels/Llama4/01-Llama4-对话助手/01-Llama4-对话助手.md291 except Exception as e:CODE
MEDIUMmodels/Llama4/01-Llama4-对话助手/01-Llama4-对话助手.md217def generate():CODE
MEDIUMmodels/Llama4/01-Llama4-对话助手/01-Llama4-对话助手.md281def clear_history():CODE
LOWmodels/Llama4/01-Llama4-对话助手/app/app.py129 except Exception as e:CODE
LOWmodels/Llama4/01-Llama4-对话助手/app/app.py143 except Exception as e:CODE
MEDIUMmodels/Llama4/01-Llama4-对话助手/app/app.py69def generate():CODE
MEDIUMmodels/Llama4/01-Llama4-对话助手/app/app.py133def clear_history():CODE
LOWmodels/Qwen3-VL/02-Qwen3-VL-4B-Instruct FastApi 部署调用.md188 except Exception as e:CODE
LOWmodels/Qwen3-VL/02-Qwen3-VL-4B-Instruct FastApi 部署调用.md255 except Exception as e:CODE
LOWmodels/Qwen3-VL/02-Qwen3-VL-4B-Instruct FastApi 部署调用.md294 except Exception as e:CODE
LOWmodels/Qwen3-VL/02-Qwen3-VL-4B-Instruct FastApi 部署调用.md345 except Exception as e:CODE
LOWmodels/Qwen3-VL/02-Qwen3-VL-4B-Instruct FastApi 部署调用.md535 except Exception as e:CODE
LOWmodels/Qwen3-VL/02-Qwen3-VL-4B-Instruct FastApi 部署调用.md599 except Exception as e:CODE
LOWmodels/Qwen3-VL/02-Qwen3-VL-4B-Instruct FastApi 部署调用.md650 except Exception as e:CODE
LOW…VL/05-Qwen3-VL-4B-Instruct Lora 可视化微调案例 - LaTexOCR.md725 except Exception:CODE
LOW…VL/05-Qwen3-VL-4B-Instruct Lora 可视化微调案例 - LaTexOCR.md824 except Exception:CODE
LOW…VL/05-Qwen3-VL-4B-Instruct Lora 可视化微调案例 - LaTexOCR.md836 except Exception:CODE
LOW…VL/05-Qwen3-VL-4B-Instruct Lora 可视化微调案例 - LaTexOCR.md838 except Exception as e:CODE
LOW…n3-VL/Qwen3-VL-4B-Instruct Lora 可视化微调案例 - LaTexOCR.md706 except Exception:CODE
LOW…n3-VL/Qwen3-VL-4B-Instruct Lora 可视化微调案例 - LaTexOCR.md805 except Exception:CODE
LOW…n3-VL/Qwen3-VL-4B-Instruct Lora 可视化微调案例 - LaTexOCR.md817 except Exception:CODE
LOW…n3-VL/Qwen3-VL-4B-Instruct Lora 可视化微调案例 - LaTexOCR.md819 except Exception as e:CODE
LOW…L-4B-Instruct FastApi 参考代码/api_server_qwen3vl_video.py126 except Exception as e:CODE
LOW…1-Qwen3-VL-4B-Instruct FastApi 参考代码/test_simple_api.py25 except Exception as e:CODE
LOW…1-Qwen3-VL-4B-Instruct FastApi 参考代码/test_simple_api.py64 except Exception as e:CODE
LOW…1-Qwen3-VL-4B-Instruct FastApi 参考代码/test_simple_api.py115 except Exception as e:CODE
LOW…01-Qwen3-VL-4B-Instruct FastApi 参考代码/test_video_api.py28 except Exception as e:CODE
LOW…01-Qwen3-VL-4B-Instruct FastApi 参考代码/test_video_api.py79 except Exception as e:CODE
LOW…-4B-Instruct FastApi 参考代码/api_server_qwen3vl_simple.py125 except Exception as e:CODE
LOW…uct Lora 可视化微调案例 - LaTexOCR/compare_qwen3_vl_infer.py63 except Exception:CODE
LOW…uct Lora 可视化微调案例 - LaTexOCR/compare_qwen3_vl_infer.py162 except Exception:CODE
LOW…uct Lora 可视化微调案例 - LaTexOCR/compare_qwen3_vl_infer.py174 except Exception:CODE
LOW…uct Lora 可视化微调案例 - LaTexOCR/compare_qwen3_vl_infer.py176 except Exception as e:CODE
LOWmodels/gpt-oss/4-gpt-oss-20b Lora 微调及 SwanLab 可视化记录.md245except Exception as e:CODE
LOWmodels/gpt-oss/3-gpt-oss-20b lmstudio 本地部署调用.md130 except Exception as exc:CODE
LOWmodels/ChatGLM/05-ChatGLM3-6B接入LangChain搭建知识库助手.md398 except Exception as e:CODE
LOWmodels/ChatGLM/Extra-ChatGLM3-6B-Pipeline.md341except Exception as e:CODE
LOWmodels/ChatGLM/Extra-ChatGLM3-6B-Pipeline.md404 except Exception as e:CODE
LOWmodels/ChatGLM/Extra-ChatGLM3-6B-Pipeline.md457 except Exception as e:CODE
LOWmodels/ChatGLM/Extra-ChatGLM3-6B-Pipeline.md491 except Exception as e:CODE
LOW…ChatGLM/05-ChatGLM3-6B接入LangChain搭建知识库助手/run_gradio.py71 except Exception as e:CODE
LOWmodels/InternLM/06-InternLM接入LangChain搭建知识库助手.md401 except Exception as e:CODE
LOW…s/InternLM/06-InternLM接入LangChain搭建知识库助手/run_gradio.py61 except Exception as e:CODE
LOW…VL/02-Qwen2-VL-2B-Instruct Web Demo 参考代码/mm_qwen2vl.py167 except Exception:CODE
LOWmodels/Qwen2.5/06-Qwen2.5-7B-Instruct o1-like 推理链实现.md118 except Exception as e:CODE
LOWmodels/Gemma3/8-gemma3-4b-it 模型服务部署.md86 except Exception as e:CODE
LOWmodels/Gemma3/01-gemma-3-4b-it FastApi 部署调用.md136 except Exception as e:CODE
LOWmodels/Gemma3/01-gemma-3-4b-it FastApi 部署调用.md146 except Exception as e:CODE
LOWmodels/Gemma3/01-gemma-3-4b-it FastApi 部署调用.md210 except Exception as e:CODE
LOWmodels/Atom/02-Atom-7B-Chat Lora 微调.md298except Exception as e:CODE
LOWmodels/Atom/03-Atom-7B-Chat 接入langchain搭建知识库助手.md443 except Exception as e:CODE
52 more matches not shown…
Decorative Section Separators16 hits · 57 pts
SeverityFileLineSnippetContext
MEDIUMmodels/Qwen3/10-Qwen3-8B GRPO微调及通过swanlab可视化.md427# ========================COMMENT
MEDIUMmodels/Qwen3/10-Qwen3-8B GRPO微调及通过swanlab可视化.md429# ========================COMMENT
MEDIUMmodels_mlx/run_app_gradio.py20# ============================================================COMMENT
MEDIUMmodels_mlx/run_app_gradio.py22# ============================================================COMMENT
MEDIUMmodels_mlx/run_app_gradio.py131# ============================================================COMMENT
MEDIUMmodels_mlx/run_app_gradio.py133# ============================================================COMMENT
MEDIUMmodels_mlx/run_app_gradio.py307# ============================================================COMMENT
MEDIUMmodels_mlx/run_app_gradio.py309# ============================================================COMMENT
MEDIUMmodels_mlx/modules/framework.py104# ============================================================COMMENT
MEDIUMmodels_mlx/modules/framework.py106# ============================================================COMMENT
MEDIUMmodels_mlx/modules/download_model.py14# ============================================================COMMENT
MEDIUMmodels_mlx/modules/download_model.py16# ============================================================COMMENT
MEDIUMmodels_mlx/modules/download_model.py23# ============================================================COMMENT
MEDIUMmodels_mlx/modules/download_model.py25# ============================================================COMMENT
MEDIUMmodels_mlx/modules/download_model.py183# ============================================================COMMENT
MEDIUMmodels_mlx/modules/download_model.py185# ============================================================COMMENT
Unused Imports34 hits · 34 pts
SeverityFileLineSnippetContext
LOWutils.py2CODE
LOWutils.py3CODE
LOW…uct Lora 可视化微调案例 - LaTexOCR/compare_qwen3_vl_infer.py2CODE
LOW…uct Lora 可视化微调案例 - LaTexOCR/compare_qwen3_vl_infer.py3CODE
LOW…uct Lora 可视化微调案例 - LaTexOCR/compare_qwen3_vl_infer.py3CODE
LOW…ChatGLM/05-ChatGLM3-6B接入LangChain搭建知识库助手/run_gradio.py15CODE
LOW…s/InternLM/06-InternLM接入LangChain搭建知识库助手/run_gradio.py5CODE
LOW…2B-Instruct FastApi 参考代码/api_server_image_and_video.py3CODE
LOW…-Qwen2-VL-2B-Instruct FastApi 参考代码/api_server_image.py3CODE
LOW…-Instruct FastApi 参考代码/qwen_vl_utils/vision_process.py1CODE
LOW…/Atom/03-Atom-7B-Chat 接入langchain搭建知识库助手/run_gradio.py5CODE
LOWmodels/Atom/02-Atom-7B-Chat-Lora/train.py15CODE
LOWmodels/Atom/02-Atom-7B-Chat-Lora/train.py18CODE
LOWmodels/Gemma4/api.py2CODE
LOW…dels/MiniCPM-o/04-MiniCPM-0-2.6 Lora微调 参考代码/trainer.py5CODE
LOW…CPM-o/04-MiniCPM-0-2.6 Lora微调 参考代码/minicpm_datasets.py2CODE
LOW…CPM-o/04-MiniCPM-0-2.6 Lora微调 参考代码/minicpm_datasets.py5CODE
LOW…CPM-o/04-MiniCPM-0-2.6 Lora微调 参考代码/minicpm_datasets.py8CODE
LOW…CPM-o/04-MiniCPM-0-2.6 Lora微调 参考代码/minicpm_datasets.py8CODE
LOW…CPM-o/04-MiniCPM-0-2.6 Lora微调 参考代码/minicpm_datasets.py9CODE
LOW…CPM-o/04-MiniCPM-0-2.6 Lora微调 参考代码/minicpm_datasets.py9CODE
LOW…CPM-o/04-MiniCPM-0-2.6 Lora微调 参考代码/minicpm_datasets.py16CODE
LOW…CPM-o/04-MiniCPM-0-2.6 Lora微调 参考代码/minicpm_datasets.py16CODE
LOWmodels/MiniCPM/train.py6CODE
LOWmodels/Kimi-VL/01-Kimi-VL-对话助手/app/app.py9CODE
LOW…/Qwen/07-Qwen-7B-Chat 接入langchain搭建知识库助手/run_gradio.py5CODE
LOWmodels/XVERSE/code/LLM.py4CODE
LOWmodels/XVERSE/code/LLM.py4CODE
LOWmodels/XVERSE/code/model_download.py1CODE
LOWmodels/XVERSE/code/model_download.py2CODE
LOWmodels/XVERSE/code/model_download.py2CODE
LOWmodels/XVERSE/code/model_download.py3CODE
LOWmodels/BlueLM/04-BlueLM-7B-Chat Lora 微调.py3CODE
LOWexamples/Chat-嬛嬛/train.py4CODE
Deep Nesting19 hits · 19 pts
SeverityFileLineSnippetContext
LOWutils.py91CODE
LOW…/ChatGLM/05-ChatGLM3-6B接入LangChain搭建知识库助手/create_db.py23CODE
LOW…els/InternLM/06-InternLM接入LangChain搭建知识库助手/creat_db.py11CODE
LOW…-Instruct FastApi 参考代码/qwen_vl_utils/vision_process.py82CODE
LOW…-Instruct FastApi 参考代码/qwen_vl_utils/vision_process.py303CODE
LOW…VL/02-Qwen2-VL-2B-Instruct Web Demo 参考代码/mm_qwen2vl.py54CODE
LOW…VL/02-Qwen2-VL-2B-Instruct Web Demo 参考代码/mm_qwen2vl.py229CODE
LOW…ls/Atom/03-Atom-7B-Chat 接入langchain搭建知识库助手/creat_db.py13CODE
LOWmodels/Gemma4/verify_gemma4_tutorials.py66CODE
LOWmodels/Gemma4/verify_gemma4_tutorials.py134CODE
LOWmodels/Gemma4/api.py107CODE
LOW…CPM-o/04-MiniCPM-0-2.6 Lora微调 参考代码/minicpm_datasets.py231CODE
LOW…CPM-o/04-MiniCPM-0-2.6 Lora微调 参考代码/minicpm_datasets.py271CODE
LOW…CPM-o/04-MiniCPM-0-2.6 Lora微调 参考代码/minicpm_datasets.py310CODE
LOW…CPM-o/04-MiniCPM-0-2.6 Lora微调 参考代码/minicpm_datasets.py427CODE
LOWmodels/Kimi-VL/01-Kimi-VL-对话助手/app/app.py167CODE
LOW…ls/Qwen/07-Qwen-7B-Chat 接入langchain搭建知识库助手/creat_db.py11CODE
LOWmodels_mlx/run_app_gradio.py329CODE
LOWmodels_mlx/modules/download_model.py120CODE
Structural Annotation Overuse11 hits · 16 pts
SeverityFileLineSnippetContext
LOWmodels/Qwen3/10-Qwen3-8B GRPO微调及通过swanlab可视化.md717# Step 1: 构造初始输入并生成输出(不加载 LoRA)COMMENT
LOWmodels/Qwen3/10-Qwen3-8B GRPO微调及通过swanlab可视化.md738# Step 2: 保存 GRPO 微调得到的 LoRA 权重COMMENT
LOWmodels/Qwen3/10-Qwen3-8B GRPO微调及通过swanlab可视化.md741# Step 3: 检查保存的 safetensors 权重不为全零COMMENT
LOWmodels/Qwen3/10-Qwen3-8B GRPO微调及通过swanlab可视化.md752# Step 4: 构造消息格式输入并应用 tokenizer 的 chat_templateCOMMENT
LOWmodels/Qwen3/10-Qwen3-8B GRPO微调及通过swanlab可视化.md765# Step 5: 加载微调后的 LoRA 并生成输出COMMENT
LOWexamples/Chat-嬛嬛/readme.md15## Step 1: 环境准备COMMENT
LOWexamples/Chat-嬛嬛/readme.md44## Step 2: 数据准备COMMENT
LOWexamples/Chat-嬛嬛/readme.md130## Step 3: 模型训练COMMENT
LOWexamples/AMchat-高等数学/readme.md26### Step 1: 数据准备COMMENT
LOWexamples/AMchat-高等数学/readme.md49### Step 2: 环境准备COMMENT
LOWexamples/AMchat-高等数学/readme.md66### Step 3: 模型微调COMMENT
Hyper-Verbose Identifiers13 hits · 12 pts
SeverityFileLineSnippetContext
LOWmodels/Gemma3/6-gemma3-4B-itGRPO微调及通过swanlab可视化.md258def match_format_approximately(completions, **kwargs):CODE
LOWmodels/Gemma4/verify_gemma4_tutorials.py66def test_fastapi_smoke_with_mock() -> bool:CODE
LOWmodels/Gemma4/6-gemma4-E4B-itGRPO微调及通过swanlab可视化.md260def match_format_approximately(completions, **kwargs):CODE
LOWmodels/phi4/06-Phi-4-GRPO及swanlab可视化.md229def strict_format_reward_func(completions, **kwargs) -> list[float]:STRING
LOWmodels/MiniCPM-o/04-MiniCPM-0-2.6 Lora微调.md259def make_supervised_data_module(CODE
LOWmodels/MiniCPM-o/04-MiniCPM-0-2.6 Lora微调 参考代码/train.py17def make_supervised_data_module(CODE
LOW…CPM-o/04-MiniCPM-0-2.6 Lora微调 参考代码/minicpm_datasets.py199def conversation_to_ids_minicpm(conversation, tokenizer):CODE
LOWmodels/Qwen3/10-Qwen3-8B GRPO微调及通过swanlab可视化.md474def match_format_approximately(completions, **kwargs):CODE
LOW…n-A13B-Instruct/03-Hunyuan-A13B-Instruct-SGLang部署调用.md157def chatcompletionwiththinkbyargs(messages, enable_thinking=True):CODE
LOW…l-Qwen/05-DeepSeek-R1-0528-Qwen3-8B-GRPO及swanlab可视化.md198def match_format_approximately(completions, **kwargs):CODE
LOW…l-Qwen/05-DeepSeek-R1-0528-Qwen3-8B-GRPO及swanlab可视化.md287def format_and_language_reward_func(completions, **kwargs):CODE
LOWexamples/数字生命/readme.md23def convert_to_sharegpt_format(original_data,new_system_value=None):CODE
LOWexamples/Tianji-天机/readme.md225def extract_and_merge_conversations(folder_path, output_file):STRING
Over-Commented Block11 hits · 11 pts
SeverityFileLineSnippetContext
LOWmodels/GLM-4/04-GLM-4-9B-Chat vLLM 部署调用.md101 # tokenizer = AutoTokenizer.from_pretrained(model, use_fast=False) # 加载分词器后传入vLLM 模型,但不是必要的。COMMENT
LOWmodels/Qwen2.5/03-Qwen2.5-7B-Instruct vLLM 部署调用.md101COMMENT
LOWmodels/MiniCPM-o/03-MiniCPM-o-2.6 多模态语音能力.md141COMMENT
LOWmodels/Qwen2/04-Qwen2-7B-Instruct vLLM 部署调用.md121 # tokenizer = AutoTokenizer.from_pretrained(model, use_fast=False) COMMENT
LOWmodels/Qwen/08-Qwen-7B-Chat Lora 低精度微调.py61 tokenized_id = ds.map(process_func, remove_columns=ds.column_names)COMMENT
LOWmodels/Qwen1.5/07-Qwen1.5-7B-Chat vLLM 推理部署调用.md101if __name__ == "__main__": COMMENT
LOWmodels/XVERSE/code/data_format.py21with open('ruozhiba.json', 'w', encoding='utf-8') as f:COMMENT
LOWmodels/InternLM3/04-InternLM3-8B-Instruct LoRA.md201> 注意:此处要记得修改为自己的模型路径哦~COMMENT
LOWmodels/InternLM3/04-InternLM3-8B-Instruct LoRA.md221# )COMMENT
LOW…still-Qwen/04-DeepSeek-R1-Distill-Qwen-7B vLLM 部署调用.md101 # tokenizer = AutoTokenizer.from_pretrained(model, use_fast=False) COMMENT
LOWexamples/Tianji-天机/readme.md41from zhipuai import ZhipuAICOMMENT
Redundant / Tautological Comments7 hits · 10 pts
SeverityFileLineSnippetContext
LOWutils.py21 # Read filesCOMMENT
LOWutils.py61 # Check if the task contains "Lora" or "微调" (case-insensitive)COMMENT
LOWutils.py84 # Print resultsCOMMENT
LOWmodels/GLM-4/benchmark_throughput.py155 # Check if we can add more requests to the batch.COMMENT
LOWmodels/Qwen2.5/benchmark_throughput.py257 # Check if we can add more requests to the batch.COMMENT
LOWmodels/Qwen2/benchmark_throughput.py154 # Check if we can add more requests to the batch.COMMENT
LOWmodels/Qwen1.5/benchmark_throughput.py154 # Check if we can add more requests to the batch.COMMENT
Docstring Block Structure1 hit · 5 pts
SeverityFileLineSnippetContext
HIGH…-Instruct FastApi 参考代码/qwen_vl_utils/vision_process.py132calculate the number of frames for video used for model inputs. Args: ele (dict): a dict contains the confiSTRING
Magic Placeholder Names1 hit · 5 pts
SeverityFileLineSnippetContext
HIGH…dels/GLM-4.7-Flash/03-GLM-4.7-Flash-Lora微调及Docker镜像.md137swanlab.login(api_key='your-apikey', save=True) # 记得替换为自己账号的apikeyCODE
Modern Structural Boilerplate3 hits · 3 pts
SeverityFileLineSnippetContext
LOW…-Instruct FastApi 参考代码/qwen_vl_utils/vision_process.py22logger = logging.getLogger(__name__)CODE
LOWmodels/Gemma4/api.py17logger = logging.getLogger(__name__)CODE
LOW…CPM-o/04-MiniCPM-0-2.6 Lora微调 参考代码/minicpm_datasets.py19logger = logging.getLogger(__name__)CODE
Example Usage Blocks1 hit · 2 pts
SeverityFileLineSnippetContext
LOWutils.py149# Usage exampleCOMMENT
AI Structural Patterns1 hit · 1 pts
SeverityFileLineSnippetContext
LOWmodels/XVERSE/code/LLM.py30CODE