初始化红餐观察库项目
This commit is contained in:
@@ -0,0 +1,4 @@
|
||||
DASHSCOPE_API_KEY=replace_with_your_key
|
||||
QWEN_MODEL=qwen-plus
|
||||
FLASK_SECRET_KEY=replace_with_a_random_local_secret
|
||||
PORT=8766
|
||||
@@ -0,0 +1,7 @@
|
||||
.env
|
||||
.venv/
|
||||
__pycache__/
|
||||
*.pyc
|
||||
data/*.db-shm
|
||||
data/*.db-wal
|
||||
|
||||
@@ -0,0 +1 @@
|
||||
[ 634ms] [ERROR] Failed to load resource: the server responded with a status of 404 (NOT FOUND) @ http://127.0.0.1:8766/favicon.ico:0
|
||||
@@ -0,0 +1,189 @@
|
||||
- generic [active] [ref=e1]:
|
||||
- banner [ref=e2]:
|
||||
- link "返回文章库首页" [ref=e3] [cursor=pointer]:
|
||||
- /url: /
|
||||
- generic [ref=e4]: HC
|
||||
- generic [ref=e5]:
|
||||
- generic [ref=e6]: 红餐观察库
|
||||
- generic [ref=e7]: ARTICLE INTELLIGENCE
|
||||
- navigation [ref=e8]:
|
||||
- link "文章库" [ref=e9] [cursor=pointer]:
|
||||
- /url: /
|
||||
- generic [ref=e11]: 本地运行
|
||||
- main [ref=e12]:
|
||||
- generic [ref=e13]:
|
||||
- generic [ref=e14]:
|
||||
- paragraph [ref=e15]: 餐饮行业内容情报
|
||||
- heading "从文章堆里, 读出行业信号。" [level=1] [ref=e16]: 从文章堆里,读出行业信号。
|
||||
- search [ref=e17]:
|
||||
- generic [ref=e18]: 搜索标题、品牌或正文
|
||||
- generic [ref=e19]:
|
||||
- textbox "搜索标题、品牌或正文" [ref=e20]:
|
||||
- /placeholder: 例如:咖啡、海底捞、供应链
|
||||
- button "搜索" [ref=e21] [cursor=pointer]
|
||||
- paragraph [ref=e22]: 已收录 129 篇,1 篇完成 AI 分析
|
||||
- region "文章库数据" [ref=e23]:
|
||||
- generic [ref=e24]:
|
||||
- strong [ref=e25]: "129"
|
||||
- generic [ref=e26]: 完整文章
|
||||
- generic [ref=e27]:
|
||||
- strong [ref=e28]: "1"
|
||||
- generic [ref=e29]: AI 分析
|
||||
- generic [ref=e30]:
|
||||
- strong [ref=e31]: 2026.07.19 11:24
|
||||
- generic [ref=e32]: 最近更新
|
||||
- generic [ref=e34]:
|
||||
- heading "最新入库" [level=2] [ref=e35]
|
||||
- paragraph [ref=e36]: 共 129 篇
|
||||
- generic [ref=e37]:
|
||||
- article [ref=e38]:
|
||||
- time [ref=e40]: 2026.07.19 11:24
|
||||
- heading [level=3] [ref=e41]:
|
||||
- link "重庆市烹饪协会第八届第一次会员代表大会圆满召开" [ref=e42] [cursor=pointer]:
|
||||
- /url: /article/30
|
||||
- paragraph [ref=e43]: 重庆市烹饪协会第八届第一次会员代表大会在重庆金陵大饭店隆重召开。
|
||||
- link "阅读全文 ↗" [ref=e44] [cursor=pointer]:
|
||||
- /url: /article/30
|
||||
- article [ref=e45]:
|
||||
- time [ref=e47]: 2026.07.17 18:29
|
||||
- heading [level=3] [ref=e48]:
|
||||
- link "自己种的食材自己开餐厅卖,餐饮业迎来一批“农科所”餐厅" [ref=e49] [cursor=pointer]:
|
||||
- /url: /article/1
|
||||
- paragraph [ref=e50]: 采摘鲜蔬直接下锅,从田间到餐桌零延迟,这般极致新鲜的火锅体验,正俘获大批食客。
|
||||
- link "阅读全文 ↗" [ref=e51] [cursor=pointer]:
|
||||
- /url: /article/1
|
||||
- article [ref=e52]:
|
||||
- time [ref=e54]: 2026.07.17 16:38
|
||||
- heading [level=3] [ref=e55]:
|
||||
- link "圆桌对话:浪潮转变,坚韧生长——新全球化周期中的突破与共生" [ref=e56] [cursor=pointer]:
|
||||
- /url: /article/2
|
||||
- paragraph [ref=e57]: “看国内商场,看到的都是竞争;看海外,都是蓝海。”
|
||||
- link "阅读全文 ↗" [ref=e58] [cursor=pointer]:
|
||||
- /url: /article/2
|
||||
- article [ref=e59]:
|
||||
- time [ref=e61]: 2026.07.17 16:36
|
||||
- heading [level=3] [ref=e62]:
|
||||
- link "一碗小面、一笼包点里的世界杯:抖音地方特色早餐订单最高涨212%" [ref=e63] [cursor=pointer]:
|
||||
- /url: /article/3
|
||||
- paragraph [ref=e64]: 随着美加墨世界杯进入收官阶段,赛事热度正从线上讨论持续延伸至线下消费。
|
||||
- link "阅读全文 ↗" [ref=e65] [cursor=pointer]:
|
||||
- /url: /article/3
|
||||
- article [ref=e66]:
|
||||
- time [ref=e68]: 2026.07.17 09:02
|
||||
- heading [level=3] [ref=e69]:
|
||||
- link "山东老牌酒楼正式闭店!耗资过亿元打造,曾火到一桌难求" [ref=e70] [cursor=pointer]:
|
||||
- /url: /article/4
|
||||
- paragraph [ref=e71]: 曾经一天同时6场婚礼、订席面要提前大半年的酒店,如今大门紧闭,只剩一张红纸黑字的停业告示。
|
||||
- link "阅读全文 ↗" [ref=e72] [cursor=pointer]:
|
||||
- /url: /article/4
|
||||
- article [ref=e73]:
|
||||
- time [ref=e75]: 2026.07.17 08:59
|
||||
- heading [level=3] [ref=e76]:
|
||||
- link "幸运咖计划下半年新增门店控制在1000家以内;海底捞推出汉堡品牌“欢鲜堡”" [ref=e77] [cursor=pointer]:
|
||||
- /url: /article/5
|
||||
- paragraph [ref=e78]: 幸运咖计划下半年新增门店控制在1000家以内;海底捞推出“欢鲜堡”。详情请看红餐网《每日餐讯》。
|
||||
- link "阅读全文 ↗" [ref=e79] [cursor=pointer]:
|
||||
- /url: /article/5
|
||||
- article [ref=e80]:
|
||||
- time [ref=e82]: 2026.07.16 17:14
|
||||
- heading [level=3] [ref=e83]:
|
||||
- link "加码咖啡,发力下午茶!华莱士推出7款果咖,最低5.9元起" [ref=e84] [cursor=pointer]:
|
||||
- /url: /article/7
|
||||
- paragraph [ref=e85]: 上新7款风味果咖的同时,还同步推出多款下午茶优惠套餐,华莱士要在咖啡赛道上持续加码?
|
||||
- link "阅读全文 ↗" [ref=e86] [cursor=pointer]:
|
||||
- /url: /article/7
|
||||
- article [ref=e87]:
|
||||
- time [ref=e89]: 2026.07.16 17:06
|
||||
- heading [level=3] [ref=e90]:
|
||||
- link "一张配方引发的争议:厨师离职该不该把秘方留给老板?" [ref=e91] [cursor=pointer]:
|
||||
- /url: /article/6
|
||||
- paragraph [ref=e92]: 厨师离职,老板竟然要求无偿交出配方,这合理吗?
|
||||
- link "阅读全文 ↗" [ref=e93] [cursor=pointer]:
|
||||
- /url: /article/6
|
||||
- article [ref=e94]:
|
||||
- time [ref=e96]: 2026.07.16 16:27
|
||||
- heading [level=3] [ref=e97]:
|
||||
- link "蜀海供应链南京新仓盛大开仓" [ref=e98] [cursor=pointer]:
|
||||
- /url: /article/8
|
||||
- paragraph [ref=e99]: 立体交通网络,多级配送协同。
|
||||
- link "阅读全文 ↗" [ref=e100] [cursor=pointer]:
|
||||
- /url: /article/8
|
||||
- article [ref=e101]:
|
||||
- time [ref=e103]: 2026.07.16 14:10
|
||||
- heading [level=3] [ref=e104]:
|
||||
- link "柠季×纸嫁衣 | “九世约,只待柠”:红心芭乐,清甜赴约" [ref=e105] [cursor=pointer]:
|
||||
- /url: /article/9
|
||||
- paragraph [ref=e106]: 柠季持续深耕IP跨界,用产品与情感对话年轻人。
|
||||
- link "阅读全文 ↗" [ref=e107] [cursor=pointer]:
|
||||
- /url: /article/9
|
||||
- article [ref=e108]:
|
||||
- time [ref=e110]: 2026.07.16 14:04
|
||||
- heading [level=3] [ref=e111]:
|
||||
- link "Popeyes联合淘宝闪购加速品牌战略合作,首次在中国探索AI分析和小店模型" [ref=e112] [cursor=pointer]:
|
||||
- /url: /article/10
|
||||
- paragraph [ref=e113]: Popeyes宣布与淘宝闪购深化战略合作。
|
||||
- link "阅读全文 ↗" [ref=e114] [cursor=pointer]:
|
||||
- /url: /article/10
|
||||
- article [ref=e115]:
|
||||
- time [ref=e117]: 2026.07.16 11:31
|
||||
- heading [level=3] [ref=e118]:
|
||||
- link "茶咖茉莉花供应链告急!涨超200%,横州花价冲破50元一斤" [ref=e119] [cursor=pointer]:
|
||||
- /url: /article/11
|
||||
- paragraph [ref=e120]: 尽管洪水已经退去!但这场洪灾对茶咖行业的影响,才刚刚开始显现。
|
||||
- link "阅读全文 ↗" [ref=e121] [cursor=pointer]:
|
||||
- /url: /article/11
|
||||
- article [ref=e122]:
|
||||
- time [ref=e124]: 2026.07.16 11:15
|
||||
- heading [level=3] [ref=e125]:
|
||||
- link "从鲜橙、西瓜到蜜瓜接棒,书亦烧仙草把“大众水果”卖疯了?" [ref=e126] [cursor=pointer]:
|
||||
- /url: /article/12
|
||||
- paragraph [ref=e127]: 供应链稳定+口感普适成就“水果”爆款出圈。
|
||||
- link "阅读全文 ↗" [ref=e128] [cursor=pointer]:
|
||||
- /url: /article/12
|
||||
- article [ref=e129]:
|
||||
- time [ref=e131]: 2026.07.16 09:28
|
||||
- heading [level=3] [ref=e132]:
|
||||
- link "板前餐饮爆火背后,90%的门店都只是徒有其表的刻意作秀?" [ref=e133] [cursor=pointer]:
|
||||
- /url: /article/13
|
||||
- paragraph [ref=e134]: 如今的板前餐饮,正站在“一哄而上”的风口上。
|
||||
- link "阅读全文 ↗" [ref=e135] [cursor=pointer]:
|
||||
- /url: /article/13
|
||||
- article [ref=e136]:
|
||||
- time [ref=e138]: 2026.07.16 09:24
|
||||
- heading [level=3] [ref=e139]:
|
||||
- link "一根20年前的吸管,被瑞幸重新“捡”起来了" [ref=e140] [cursor=pointer]:
|
||||
- /url: /article/14
|
||||
- paragraph [ref=e141]: 这次瑞幸选择开外挂。
|
||||
- link "阅读全文 ↗" [ref=e142] [cursor=pointer]:
|
||||
- /url: /article/14
|
||||
- article [ref=e143]:
|
||||
- time [ref=e145]: 2026.07.16 09:21
|
||||
- heading [level=3] [ref=e146]:
|
||||
- link "供应链名品 | 鲜美来鱼皮虾滑:高颜值、低热量,可以涮火锅、做捞汁、配面条" [ref=e147] [cursor=pointer]:
|
||||
- /url: /article/15
|
||||
- paragraph [ref=e148]: 餐饮行业的品类边界正在模糊。火锅店卖甜品,茶饮店做烘焙,业态融合的背后,是对场景增量的追逐。
|
||||
- link "阅读全文 ↗" [ref=e149] [cursor=pointer]:
|
||||
- /url: /article/15
|
||||
- article [ref=e150]:
|
||||
- time [ref=e152]: 2026.07.16 09:19
|
||||
- heading [level=3] [ref=e153]:
|
||||
- link "华莱士开卖下午茶!蜀海供应链南京新仓投运" [ref=e154] [cursor=pointer]:
|
||||
- /url: /article/16
|
||||
- paragraph [ref=e155]: 8.9元起!华莱士开卖下午茶;蜀海供应链南京新仓投运,华东仓网持续加密。详情请看红餐网《每日餐讯》。
|
||||
- link "阅读全文 ↗" [ref=e156] [cursor=pointer]:
|
||||
- /url: /article/16
|
||||
- article [ref=e157]:
|
||||
- time [ref=e159]: 2026.07.16 08:58
|
||||
- heading [level=3] [ref=e160]:
|
||||
- link "2026风向标产品(秋季)全球发布会及餐饮产业供需对接会将在广州举办" [ref=e161] [cursor=pointer]:
|
||||
- /url: /article/17
|
||||
- paragraph [ref=e162]: 2026风向标产品(秋季)全球发布会与2026(秋季)“从源头到餐桌”产品发布暨餐饮产业供需对接会两大核心板块,共同打造集产品发布、供需对接、传播推介于一体的餐饮产业发布平台。
|
||||
- link "阅读全文 ↗" [ref=e163] [cursor=pointer]:
|
||||
- /url: /article/17
|
||||
- navigation "分页" [ref=e164]:
|
||||
- generic [ref=e165]: 1 / 8
|
||||
- link "下一页" [ref=e166] [cursor=pointer]:
|
||||
- /url: "?q=&page=2"
|
||||
- contentinfo [ref=e167]:
|
||||
- generic [ref=e168]: 红餐文章本地研究工具
|
||||
- generic [ref=e169]: SQLite · Playwright · Qwen
|
||||
Binary file not shown.
|
After Width: | Height: | Size: 55 KiB |
@@ -0,0 +1,72 @@
|
||||
# 红餐网文章增量入库
|
||||
|
||||
使用 Playwright 真实浏览器读取红餐网移动资讯列表,将文章增量保存到 SQLite。数据库包含文章表和采集运行日志表。
|
||||
|
||||
## 安装
|
||||
|
||||
```bash
|
||||
cd /path/to/hongcan_ingestor
|
||||
python3 -m venv .venv
|
||||
source .venv/bin/activate
|
||||
pip install -r requirements.txt
|
||||
playwright install chromium
|
||||
```
|
||||
|
||||
复制 `.env.example` 为 `.env`,填写阿里云百炼 API Key。Key 只保存在本机,不要提交到 Git。
|
||||
|
||||
## 运行
|
||||
|
||||
首次运行或手动测试:
|
||||
|
||||
```bash
|
||||
python3 hongcan_ingest.py --db ./data/hongcan.db --headed
|
||||
```
|
||||
|
||||
每日定时运行:
|
||||
|
||||
```bash
|
||||
python3 hongcan_ingest.py --db ./data/hongcan.db --lookback-days 7
|
||||
```
|
||||
|
||||
查看最近采集结果:
|
||||
|
||||
```bash
|
||||
sqlite3 ./data/hongcan.db \
|
||||
"SELECT published_at,title,canonical_url FROM articles WHERE crawl_status='success' ORDER BY published_at DESC LIMIT 20;"
|
||||
```
|
||||
|
||||
crontab 示例(每天 08:10):
|
||||
|
||||
```cron
|
||||
10 8 * * * cd /path/to/hongcan_ingestor && .venv/bin/python hongcan_ingest.py --db ./data/hongcan.db >> ./data/cron.log 2>&1
|
||||
```
|
||||
|
||||
## 增量规则
|
||||
|
||||
- `canonical_url` 唯一索引,重复运行不会重复插入。
|
||||
- 默认回溯最近 7 天,避免文章晚发布或列表加载遗漏。
|
||||
- 抓取失败最多重试 3 次,错误保存在 `last_error`。
|
||||
- 正文与标题生成 SHA-256 内容指纹。
|
||||
- 每次执行情况记录在 `crawl_runs`。
|
||||
|
||||
如果无头模式被拦截,可先使用 `--headed` 检查是否出现验证码。
|
||||
|
||||
## 启动文章库前端
|
||||
|
||||
```bash
|
||||
./run_web.sh
|
||||
```
|
||||
|
||||
浏览器打开 `http://127.0.0.1:8766`。前端支持全文关键词搜索、文章详情阅读、原文跳转和千问 AI 分析。分析结果保存在数据库的 `ai_analyses` 表;文章正文未变化时会直接使用缓存。
|
||||
|
||||
批量分析全部文章:
|
||||
|
||||
```bash
|
||||
python3 batch_ai_analyze.py
|
||||
```
|
||||
|
||||
从近期文章分析中提炼跨文章行业信号:
|
||||
|
||||
```bash
|
||||
python3 extract_industry_signals.py --days 30
|
||||
```
|
||||
@@ -0,0 +1,220 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Local Hongcan article library and Qwen-assisted editorial analysis."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import os
|
||||
import sqlite3
|
||||
import urllib.error
|
||||
import urllib.request
|
||||
from datetime import datetime, timezone
|
||||
from pathlib import Path
|
||||
|
||||
from dotenv import load_dotenv
|
||||
from flask import Flask, abort, jsonify, render_template, request
|
||||
|
||||
|
||||
BASE_DIR = Path(__file__).resolve().parent
|
||||
DB_PATH = Path(os.getenv("HONGCAN_DB", BASE_DIR / "data" / "hongcan.db"))
|
||||
load_dotenv(BASE_DIR / ".env")
|
||||
|
||||
app = Flask(__name__)
|
||||
app.secret_key = os.getenv("FLASK_SECRET_KEY", "hongcan-local-only")
|
||||
|
||||
|
||||
def now_iso() -> str:
|
||||
return datetime.now(timezone.utc).astimezone().isoformat(timespec="seconds")
|
||||
|
||||
|
||||
def db() -> sqlite3.Connection:
|
||||
conn = sqlite3.connect(DB_PATH)
|
||||
conn.row_factory = sqlite3.Row
|
||||
conn.execute("PRAGMA foreign_keys = ON")
|
||||
return conn
|
||||
|
||||
|
||||
def ensure_schema() -> None:
|
||||
with db() as conn:
|
||||
conn.executescript((BASE_DIR / "schema.sql").read_text(encoding="utf-8"))
|
||||
|
||||
|
||||
def clamp_int(value: str | None, default: int, low: int, high: int) -> int:
|
||||
try:
|
||||
return max(low, min(high, int(value or default)))
|
||||
except ValueError:
|
||||
return default
|
||||
|
||||
|
||||
TREND_CN = {"emerging": "新兴", "accelerating": "加速", "stable": "稳定", "declining": "下行"}
|
||||
|
||||
|
||||
@app.template_filter("date_cn")
|
||||
def date_cn(value: str | None) -> str:
|
||||
if not value:
|
||||
return "时间未知"
|
||||
try:
|
||||
return datetime.fromisoformat(value).strftime("%Y.%m.%d %H:%M")
|
||||
except ValueError:
|
||||
return value
|
||||
|
||||
|
||||
@app.get("/")
|
||||
def index():
|
||||
query = (request.args.get("q") or "").strip()
|
||||
page = clamp_int(request.args.get("page"), 1, 1, 10_000)
|
||||
per_page = 18
|
||||
where = "WHERE crawl_status='success'"
|
||||
params: list[object] = []
|
||||
if query:
|
||||
where += " AND (title LIKE ? OR author LIKE ? OR content LIKE ?)"
|
||||
needle = f"%{query}%"
|
||||
params.extend([needle, needle, needle])
|
||||
with db() as conn:
|
||||
total = conn.execute(f"SELECT COUNT(*) FROM articles {where}", params).fetchone()[0]
|
||||
rows = conn.execute(
|
||||
f"""
|
||||
SELECT a.id,a.title,a.author,a.published_at,a.summary,a.canonical_url,
|
||||
substr(a.content,1,180) AS excerpt,
|
||||
json_extract(x.result_json,'$.summary') AS ai_summary,
|
||||
CASE WHEN x.id IS NULL THEN 0 ELSE 1 END AS analyzed
|
||||
FROM articles a
|
||||
LEFT JOIN ai_analyses x ON x.article_id=a.id AND x.analysis_type='editorial'
|
||||
{where.replace('crawl_status', 'a.crawl_status')}
|
||||
ORDER BY a.published_at DESC,a.id DESC LIMIT ? OFFSET ?
|
||||
""",
|
||||
[*params, per_page, (page - 1) * per_page],
|
||||
).fetchall()
|
||||
stats = conn.execute(
|
||||
"""
|
||||
SELECT COUNT(*) total,
|
||||
SUM(crawl_status='success') success,
|
||||
MAX(published_at) latest,
|
||||
(SELECT COUNT(DISTINCT article_id) FROM ai_analyses) analyzed
|
||||
FROM articles
|
||||
"""
|
||||
).fetchone()
|
||||
signal_rows = conn.execute(
|
||||
"SELECT * FROM industry_signals ORDER BY signal_date DESC,confidence DESC LIMIT 8"
|
||||
).fetchall()
|
||||
signals = []
|
||||
for signal in signal_rows:
|
||||
item = dict(signal)
|
||||
item["tags"] = json.loads(item["tags_json"])
|
||||
item["implications"] = json.loads(item["implications_json"])
|
||||
item["article_ids"] = json.loads(item["article_ids_json"])
|
||||
item["trend"] = TREND_CN.get(item["trend"], item["trend"])
|
||||
signals.append(item)
|
||||
return render_template(
|
||||
"index.html", articles=rows, query=query, page=page, total=total,
|
||||
pages=max(1, (total + per_page - 1) // per_page), stats=stats, signals=signals,
|
||||
)
|
||||
|
||||
|
||||
@app.get("/signals/<int:signal_id>")
|
||||
def signal_detail(signal_id: int):
|
||||
with db() as conn:
|
||||
signal = conn.execute("SELECT * FROM industry_signals WHERE id=?", (signal_id,)).fetchone()
|
||||
if not signal:
|
||||
abort(404)
|
||||
signal = dict(signal)
|
||||
signal["trend"] = TREND_CN.get(signal["trend"], signal["trend"])
|
||||
article_ids = json.loads(signal["article_ids_json"])
|
||||
articles = []
|
||||
if article_ids:
|
||||
marks = ",".join("?" for _ in article_ids)
|
||||
articles = conn.execute(
|
||||
f"SELECT id,title,published_at FROM articles WHERE id IN ({marks}) ORDER BY published_at DESC",
|
||||
article_ids,
|
||||
).fetchall()
|
||||
return render_template(
|
||||
"signal.html", signal=signal, articles=articles,
|
||||
implications=json.loads(signal["implications_json"]), tags=json.loads(signal["tags_json"]),
|
||||
)
|
||||
|
||||
|
||||
@app.get("/article/<int:article_id>")
|
||||
def article(article_id: int):
|
||||
with db() as conn:
|
||||
row = conn.execute("SELECT * FROM articles WHERE id=?", (article_id,)).fetchone()
|
||||
if not row:
|
||||
abort(404)
|
||||
analysis = conn.execute(
|
||||
"SELECT * FROM ai_analyses WHERE article_id=? AND analysis_type='editorial' ORDER BY updated_at DESC LIMIT 1",
|
||||
(article_id,),
|
||||
).fetchone()
|
||||
parsed = json.loads(analysis["result_json"]) if analysis else None
|
||||
return render_template("article.html", article=row, analysis=parsed, analysis_row=analysis)
|
||||
|
||||
|
||||
def qwen_analyze(title: str, content: str) -> dict:
|
||||
api_key = os.getenv("DASHSCOPE_API_KEY", "").strip()
|
||||
if not api_key:
|
||||
raise RuntimeError("尚未配置 DASHSCOPE_API_KEY,请在 .env 中设置")
|
||||
model = os.getenv("QWEN_MODEL", "qwen-plus")
|
||||
prompt = f"""你是一名资深餐饮行业研究编辑。分析下面这篇文章,只返回合法 JSON,不要 Markdown 代码块。
|
||||
JSON 字段必须为:summary(120字内摘要)、key_points(3-6条字符串)、industry_signal(行业信号)、brands(品牌数组)、numbers(关键数据数组)、content_angles(3条可延展选题)、risk_notes(事实或表述风险数组)、tags(5-10个标签)、sentiment(positive/neutral/negative)。
|
||||
文章标题:{title}
|
||||
正文:{content[:24000]}"""
|
||||
payload = json.dumps({
|
||||
"model": model,
|
||||
"messages": [
|
||||
{"role": "system", "content": "你输出严谨、简洁、可供编辑部直接使用的结构化行业分析。"},
|
||||
{"role": "user", "content": prompt},
|
||||
],
|
||||
"temperature": 0.25,
|
||||
"response_format": {"type": "json_object"},
|
||||
}, ensure_ascii=False).encode("utf-8")
|
||||
req = urllib.request.Request(
|
||||
"https://dashscope.aliyuncs.com/compatible-mode/v1/chat/completions",
|
||||
data=payload,
|
||||
headers={"Authorization": f"Bearer {api_key}", "Content-Type": "application/json"},
|
||||
method="POST",
|
||||
)
|
||||
try:
|
||||
with urllib.request.urlopen(req, timeout=90) as response:
|
||||
body = json.loads(response.read().decode("utf-8"))
|
||||
except urllib.error.HTTPError as exc:
|
||||
detail = exc.read().decode("utf-8", errors="replace")[:500]
|
||||
raise RuntimeError(f"千问接口返回 {exc.code}:{detail}") from exc
|
||||
text = body["choices"][0]["message"]["content"].strip()
|
||||
if text.startswith("```"):
|
||||
text = text.strip("`").removeprefix("json").strip()
|
||||
return json.loads(text)
|
||||
|
||||
|
||||
@app.post("/api/articles/<int:article_id>/analyze")
|
||||
def analyze(article_id: int):
|
||||
force = bool((request.get_json(silent=True) or {}).get("force"))
|
||||
model = os.getenv("QWEN_MODEL", "qwen-plus")
|
||||
with db() as conn:
|
||||
row = conn.execute("SELECT * FROM articles WHERE id=?", (article_id,)).fetchone()
|
||||
if not row:
|
||||
return jsonify({"ok": False, "error": "文章不存在"}), 404
|
||||
cached = conn.execute(
|
||||
"SELECT * FROM ai_analyses WHERE article_id=? AND model=? AND analysis_type='editorial'",
|
||||
(article_id, model),
|
||||
).fetchone()
|
||||
if cached and cached["content_hash"] == row["content_hash"] and not force:
|
||||
return jsonify({"ok": True, "cached": True, "analysis": json.loads(cached["result_json"])})
|
||||
try:
|
||||
result = qwen_analyze(row["title"], row["content"])
|
||||
except Exception as exc:
|
||||
return jsonify({"ok": False, "error": str(exc)}), 502
|
||||
timestamp = now_iso()
|
||||
conn.execute(
|
||||
"""
|
||||
INSERT INTO ai_analyses(article_id,model,analysis_type,result_json,content_hash,created_at,updated_at)
|
||||
VALUES (?,?, 'editorial', ?,?,?,?)
|
||||
ON CONFLICT(article_id,model,analysis_type) DO UPDATE SET
|
||||
result_json=excluded.result_json,content_hash=excluded.content_hash,updated_at=excluded.updated_at
|
||||
""",
|
||||
(article_id, model, json.dumps(result, ensure_ascii=False), row["content_hash"], timestamp, timestamp),
|
||||
)
|
||||
conn.commit()
|
||||
return jsonify({"ok": True, "cached": False, "analysis": result})
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
ensure_schema()
|
||||
app.run(host="127.0.0.1", port=int(os.getenv("PORT", "8766")), debug=False)
|
||||
@@ -0,0 +1,83 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Batch Qwen analysis for all successfully crawled articles."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import os
|
||||
import sqlite3
|
||||
import time
|
||||
from pathlib import Path
|
||||
|
||||
from app import DB_PATH, ensure_schema, now_iso, qwen_analyze
|
||||
|
||||
|
||||
def run(force: bool, pause: float, retries: int) -> int:
|
||||
ensure_schema()
|
||||
model = os.getenv("QWEN_MODEL", "qwen-plus")
|
||||
conn = sqlite3.connect(DB_PATH)
|
||||
conn.row_factory = sqlite3.Row
|
||||
rows = conn.execute(
|
||||
"""
|
||||
SELECT a.id,a.title,a.content,a.content_hash,x.content_hash analyzed_hash
|
||||
FROM articles a
|
||||
LEFT JOIN ai_analyses x
|
||||
ON x.article_id=a.id AND x.model=? AND x.analysis_type='editorial'
|
||||
WHERE a.crawl_status='success'
|
||||
ORDER BY a.published_at DESC,a.id DESC
|
||||
""",
|
||||
(model,),
|
||||
).fetchall()
|
||||
pending = [r for r in rows if force or not r["analyzed_hash"] or r["analyzed_hash"] != r["content_hash"]]
|
||||
print(json.dumps({"total": len(rows), "pending": len(pending), "model": model}, ensure_ascii=False), flush=True)
|
||||
success = skipped = failed = 0
|
||||
for index, row in enumerate(rows, start=1):
|
||||
if not force and row["analyzed_hash"] == row["content_hash"]:
|
||||
skipped += 1
|
||||
continue
|
||||
last_error = None
|
||||
for attempt in range(1, retries + 2):
|
||||
try:
|
||||
result = qwen_analyze(row["title"], row["content"])
|
||||
timestamp = now_iso()
|
||||
conn.execute(
|
||||
"""
|
||||
INSERT INTO ai_analyses(article_id,model,analysis_type,result_json,content_hash,created_at,updated_at)
|
||||
VALUES (?,?, 'editorial', ?,?,?,?)
|
||||
ON CONFLICT(article_id,model,analysis_type) DO UPDATE SET
|
||||
result_json=excluded.result_json,content_hash=excluded.content_hash,updated_at=excluded.updated_at
|
||||
""",
|
||||
(row["id"], model, json.dumps(result, ensure_ascii=False), row["content_hash"], timestamp, timestamp),
|
||||
)
|
||||
conn.commit()
|
||||
success += 1
|
||||
print(f"[{index}/{len(rows)}] OK {row['title']}", flush=True)
|
||||
last_error = None
|
||||
break
|
||||
except Exception as exc:
|
||||
last_error = exc
|
||||
if attempt <= retries:
|
||||
wait = attempt * 2
|
||||
print(f"[{index}/{len(rows)}] RETRY {attempt}/{retries} {exc}", flush=True)
|
||||
time.sleep(wait)
|
||||
if last_error is not None:
|
||||
failed += 1
|
||||
print(f"[{index}/{len(rows)}] FAIL {row['title']}: {last_error}", flush=True)
|
||||
time.sleep(pause)
|
||||
total_analyzed = conn.execute("SELECT COUNT(DISTINCT article_id) FROM ai_analyses").fetchone()[0]
|
||||
conn.close()
|
||||
print(json.dumps({
|
||||
"success": success, "skipped": skipped, "failed": failed,
|
||||
"total_analyzed": total_analyzed,
|
||||
}, ensure_ascii=False), flush=True)
|
||||
return 0 if failed == 0 else 1
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
parser = argparse.ArgumentParser(description="批量执行红餐文章千问分析")
|
||||
parser.add_argument("--force", action="store_true", help="忽略缓存并重新分析")
|
||||
parser.add_argument("--pause", type=float, default=0.35, help="请求间隔秒数")
|
||||
parser.add_argument("--retries", type=int, default=2, help="单篇失败重试次数")
|
||||
args = parser.parse_args()
|
||||
raise SystemExit(run(args.force, args.pause, args.retries))
|
||||
Binary file not shown.
@@ -0,0 +1,106 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Extract cross-article restaurant industry signals from cached Qwen analyses."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import os
|
||||
import sqlite3
|
||||
import urllib.error
|
||||
import urllib.request
|
||||
from datetime import datetime, timedelta
|
||||
|
||||
from app import DB_PATH, ensure_schema, now_iso
|
||||
|
||||
|
||||
def call_qwen(items: list[dict], model: str) -> list[dict]:
|
||||
key = os.getenv("DASHSCOPE_API_KEY", "").strip()
|
||||
if not key:
|
||||
raise RuntimeError("尚未配置 DASHSCOPE_API_KEY")
|
||||
prompt = f"""你是餐饮产业首席分析师。根据以下多篇文章的结构化分析,识别跨文章、可验证、有经营意义的行业信号。
|
||||
只返回合法 JSON:{{"signals":[...]}}。每个 signal 必须包含:title(短标题)、summary(100-180字)、trend(emerging/accelerating/stable/declining)、confidence(0到1)、article_ids(至少2个证据文章ID;确实只有单篇强信号时可为1个)、implications(2-4条经营启示)、tags(3-6个标签)。
|
||||
不要把单一品牌新闻简单改写成行业信号;合并重复主题;最多输出8条;没有充分证据就少输出。
|
||||
输入:{json.dumps(items, ensure_ascii=False)}"""
|
||||
body = json.dumps({
|
||||
"model": model,
|
||||
"messages": [
|
||||
{"role": "system", "content": "你进行基于证据的餐饮行业趋势聚类,避免空泛结论。"},
|
||||
{"role": "user", "content": prompt},
|
||||
],
|
||||
"temperature": 0.2,
|
||||
"response_format": {"type": "json_object"},
|
||||
}, ensure_ascii=False).encode("utf-8")
|
||||
req = urllib.request.Request(
|
||||
"https://dashscope.aliyuncs.com/compatible-mode/v1/chat/completions",
|
||||
data=body,
|
||||
headers={"Authorization": f"Bearer {key}", "Content-Type": "application/json"},
|
||||
method="POST",
|
||||
)
|
||||
try:
|
||||
with urllib.request.urlopen(req, timeout=120) as response:
|
||||
result = json.loads(response.read().decode("utf-8"))
|
||||
except urllib.error.HTTPError as exc:
|
||||
raise RuntimeError(f"千问接口返回 {exc.code}:{exc.read().decode('utf-8', errors='replace')[:500]}") from exc
|
||||
text = result["choices"][0]["message"]["content"].strip()
|
||||
if text.startswith("```"):
|
||||
text = text.strip("`").removeprefix("json").strip()
|
||||
return json.loads(text).get("signals", [])
|
||||
|
||||
|
||||
def run(days: int) -> int:
|
||||
ensure_schema()
|
||||
model = os.getenv("QWEN_MODEL", "qwen-plus")
|
||||
cutoff = (datetime.now() - timedelta(days=days)).strftime("%Y-%m-%dT%H:%M:%S")
|
||||
conn = sqlite3.connect(DB_PATH)
|
||||
conn.row_factory = sqlite3.Row
|
||||
rows = conn.execute(
|
||||
"""
|
||||
SELECT a.id,a.title,a.published_at,x.result_json
|
||||
FROM articles a JOIN ai_analyses x ON x.article_id=a.id
|
||||
WHERE a.crawl_status='success' AND a.published_at>=?
|
||||
ORDER BY a.published_at DESC
|
||||
""",
|
||||
(cutoff,),
|
||||
).fetchall()
|
||||
items = []
|
||||
for row in rows:
|
||||
analysis = json.loads(row["result_json"])
|
||||
items.append({
|
||||
"article_id": row["id"], "title": row["title"], "published_at": row["published_at"],
|
||||
"summary": analysis.get("summary"), "key_points": analysis.get("key_points", []),
|
||||
"industry_signal": analysis.get("industry_signal"), "tags": analysis.get("tags", []),
|
||||
})
|
||||
if not items:
|
||||
print(json.dumps({"status": "skipped", "reason": "没有已分析文章"}, ensure_ascii=False))
|
||||
return 0
|
||||
signals = call_qwen(items, model)
|
||||
signal_date = datetime.now().astimezone().date().isoformat()
|
||||
timestamp = now_iso()
|
||||
conn.execute("DELETE FROM industry_signals WHERE signal_date=?", (signal_date,))
|
||||
for signal in signals:
|
||||
conn.execute(
|
||||
"""
|
||||
INSERT INTO industry_signals
|
||||
(signal_date,title,summary,trend,confidence,implications_json,article_ids_json,tags_json,model,created_at,updated_at)
|
||||
VALUES (?,?,?,?,?,?,?,?,?,?,?)
|
||||
""",
|
||||
(
|
||||
signal_date, signal["title"], signal["summary"], signal.get("trend", "emerging"),
|
||||
float(signal.get("confidence", 0.5)), json.dumps(signal.get("implications", []), ensure_ascii=False),
|
||||
json.dumps(signal.get("article_ids", [])), json.dumps(signal.get("tags", []), ensure_ascii=False),
|
||||
model, timestamp, timestamp,
|
||||
),
|
||||
)
|
||||
conn.commit()
|
||||
conn.close()
|
||||
print(json.dumps({"status": "success", "input_articles": len(items), "signals": len(signals), "date": signal_date}, ensure_ascii=False))
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
parser = argparse.ArgumentParser(description="提炼红餐文章跨文行业信号")
|
||||
parser.add_argument("--days", type=int, default=30, help="聚合最近多少天的文章")
|
||||
args = parser.parse_args()
|
||||
raise SystemExit(run(args.days))
|
||||
|
||||
@@ -0,0 +1,301 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Incrementally collect Hongcan articles into SQLite using a real browser."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import hashlib
|
||||
import json
|
||||
import re
|
||||
import sqlite3
|
||||
import sys
|
||||
import time
|
||||
from contextlib import closing
|
||||
from datetime import datetime, timedelta, timezone
|
||||
from pathlib import Path
|
||||
from urllib.parse import urljoin, urlsplit, urlunsplit
|
||||
|
||||
from playwright.sync_api import BrowserContext, Page, sync_playwright
|
||||
|
||||
|
||||
DEFAULT_LIST_URL = "https://m.canyin88.com/zixun/"
|
||||
ARTICLE_PATH_RE = re.compile(r"/zixun/\d{4}/\d{1,2}/\d{1,2}/\d+\.html$")
|
||||
DATE_RE = re.compile(r"(20\d{2})[-/.年](\d{1,2})[-/.月](\d{1,2})(?:日)?(?:\s+(\d{1,2}):(\d{2})(?::(\d{2}))?)?")
|
||||
|
||||
|
||||
def now_iso() -> str:
|
||||
return datetime.now(timezone.utc).astimezone().isoformat(timespec="seconds")
|
||||
|
||||
|
||||
def normalize_url(raw_url: str, base_url: str = DEFAULT_LIST_URL) -> str | None:
|
||||
absolute = urljoin(base_url, raw_url.strip())
|
||||
parts = urlsplit(absolute)
|
||||
host = parts.netloc.lower().removeprefix("m.").removeprefix("www.")
|
||||
if host != "canyin88.com" or not ARTICLE_PATH_RE.search(parts.path):
|
||||
return None
|
||||
return urlunsplit(("https", "www.canyin88.com", parts.path, "", ""))
|
||||
|
||||
|
||||
def clean_text(value: str | None) -> str:
|
||||
return re.sub(r"\s+", " ", value or "").strip().lstrip("\ufeff")
|
||||
|
||||
|
||||
def parse_date(value: str | None) -> str | None:
|
||||
if not value:
|
||||
return None
|
||||
match = DATE_RE.search(value)
|
||||
if not match:
|
||||
return None
|
||||
year, month, day, hour, minute, second = match.groups()
|
||||
dt = datetime(
|
||||
int(year), int(month), int(day), int(hour or 0), int(minute or 0), int(second or 0)
|
||||
)
|
||||
return dt.isoformat(timespec="seconds")
|
||||
|
||||
|
||||
def date_from_article_url(url: str) -> datetime | None:
|
||||
match = re.search(r"/zixun/(20\d{2})/(\d{1,2})/(\d{1,2})/", url)
|
||||
if not match:
|
||||
return None
|
||||
return datetime(*(int(part) for part in match.groups()))
|
||||
|
||||
|
||||
def init_db(conn: sqlite3.Connection, schema_path: Path) -> None:
|
||||
conn.executescript(schema_path.read_text(encoding="utf-8"))
|
||||
|
||||
|
||||
def discover_urls(page: Page, list_url: str, scrolls: int, pause_ms: int) -> list[str]:
|
||||
page.goto(list_url, wait_until="domcontentloaded", timeout=60_000)
|
||||
page.wait_for_timeout(2_000)
|
||||
found: set[str] = set()
|
||||
stagnant = 0
|
||||
for _ in range(scrolls + 1):
|
||||
hrefs = page.locator("a[href]").evaluate_all("els => els.map(e => e.href)")
|
||||
before = len(found)
|
||||
for href in hrefs:
|
||||
normalized = normalize_url(href, list_url)
|
||||
if normalized:
|
||||
found.add(normalized)
|
||||
stagnant = stagnant + 1 if len(found) == before else 0
|
||||
if stagnant >= 3:
|
||||
break
|
||||
page.evaluate("window.scrollTo(0, document.body.scrollHeight)")
|
||||
page.wait_for_timeout(pause_ms)
|
||||
return sorted(found, reverse=True)
|
||||
|
||||
|
||||
def meta(page: Page, selector: str) -> str:
|
||||
locator = page.locator(selector).first
|
||||
return clean_text(locator.get_attribute("content")) if locator.count() else ""
|
||||
|
||||
|
||||
def first_text(page: Page, selectors: list[str]) -> str:
|
||||
for selector in selectors:
|
||||
locator = page.locator(selector).first
|
||||
if locator.count():
|
||||
text = clean_text(locator.inner_text(timeout=3_000))
|
||||
if text:
|
||||
return text
|
||||
return ""
|
||||
|
||||
|
||||
def extract_article(page: Page, url: str) -> dict[str, str | None]:
|
||||
page.goto(url, wait_until="domcontentloaded", timeout=60_000)
|
||||
page.wait_for_timeout(800)
|
||||
document_title = clean_text(page.title())
|
||||
document_title = re.sub(r"(?:[_-]红餐网.*)$", "", document_title).strip()
|
||||
title = meta(page, 'meta[property="og:title"]') or document_title
|
||||
if not title or title == "相关推荐":
|
||||
title = first_text(page, [".title", "h1"])
|
||||
summary = (
|
||||
meta(page, 'meta[name="description"]')
|
||||
or meta(page, 'meta[property="og:description"]')
|
||||
)
|
||||
author = (
|
||||
meta(page, 'meta[name="author"]')
|
||||
or first_text(page, [".author", ".source", "[class*=author]"])
|
||||
)
|
||||
published_raw = (
|
||||
meta(page, 'meta[property="article:published_time"]')
|
||||
or first_text(page, ["time", ".time", "[class*=time]", "[class*=date]"])
|
||||
or page.locator("body").inner_text(timeout=5_000)[:1000]
|
||||
)
|
||||
content = first_text(
|
||||
page,
|
||||
[
|
||||
"article",
|
||||
".article-content",
|
||||
".article_content",
|
||||
".content",
|
||||
"[class*=detail-content]",
|
||||
"[class*=article] [class*=content]",
|
||||
],
|
||||
)
|
||||
if not content:
|
||||
paragraphs = page.locator("p").all_inner_texts()
|
||||
content = "\n\n".join(clean_text(p) for p in paragraphs if len(clean_text(p)) >= 20)
|
||||
if not title or len(content) < 80:
|
||||
raise ValueError(f"article extraction incomplete: title={bool(title)}, content_length={len(content)}")
|
||||
published_at = parse_date(published_raw)
|
||||
url_date = date_from_article_url(url)
|
||||
if url_date and (
|
||||
not published_at or datetime.fromisoformat(published_at).date() != url_date.date()
|
||||
):
|
||||
published_at = url_date.isoformat(timespec="seconds")
|
||||
content_hash = hashlib.sha256(f"{title}\n{content}".encode("utf-8")).hexdigest()
|
||||
return {
|
||||
"canonical_url": url,
|
||||
"title": title,
|
||||
"author": author or None,
|
||||
"published_at": published_at,
|
||||
"summary": summary or None,
|
||||
"content": content,
|
||||
"content_hash": content_hash,
|
||||
}
|
||||
|
||||
|
||||
def pending_urls(conn: sqlite3.Connection, discovered: list[str], max_retries: int) -> list[str]:
|
||||
rows = conn.execute(
|
||||
"SELECT canonical_url FROM articles WHERE crawl_status != 'success' AND retry_count < ?",
|
||||
(max_retries,),
|
||||
).fetchall()
|
||||
return list(dict.fromkeys(discovered + [row[0] for row in rows]))
|
||||
|
||||
|
||||
def insert_discoveries(conn: sqlite3.Connection, urls: list[str]) -> int:
|
||||
timestamp = now_iso()
|
||||
inserted = 0
|
||||
for url in urls:
|
||||
cursor = conn.execute(
|
||||
"""
|
||||
INSERT OR IGNORE INTO articles
|
||||
(canonical_url, discovered_at, updated_at)
|
||||
VALUES (?, ?, ?)
|
||||
""",
|
||||
(url, timestamp, timestamp),
|
||||
)
|
||||
inserted += cursor.rowcount
|
||||
conn.commit()
|
||||
return inserted
|
||||
|
||||
|
||||
def save_success(conn: sqlite3.Connection, article: dict[str, str | None]) -> None:
|
||||
timestamp = now_iso()
|
||||
conn.execute(
|
||||
"""
|
||||
UPDATE articles SET
|
||||
title = ?, author = ?, published_at = ?, summary = ?, content = ?,
|
||||
content_hash = ?, crawl_status = 'success', last_error = NULL,
|
||||
crawled_at = ?, updated_at = ?
|
||||
WHERE canonical_url = ?
|
||||
""",
|
||||
(
|
||||
article["title"], article["author"], article["published_at"], article["summary"],
|
||||
article["content"], article["content_hash"], timestamp, timestamp,
|
||||
article["canonical_url"],
|
||||
),
|
||||
)
|
||||
conn.commit()
|
||||
|
||||
|
||||
def save_failure(conn: sqlite3.Connection, url: str, error: Exception) -> None:
|
||||
conn.execute(
|
||||
"""
|
||||
UPDATE articles SET
|
||||
crawl_status = 'failed', retry_count = retry_count + 1,
|
||||
last_error = ?, updated_at = ?
|
||||
WHERE canonical_url = ?
|
||||
""",
|
||||
(str(error)[:1000], now_iso(), url),
|
||||
)
|
||||
conn.commit()
|
||||
|
||||
|
||||
def new_context(playwright, headless: bool) -> tuple[BrowserContext, Page]:
|
||||
browser = playwright.chromium.launch(headless=headless)
|
||||
context = browser.new_context(
|
||||
locale="zh-CN",
|
||||
timezone_id="Asia/Shanghai",
|
||||
viewport={"width": 1280, "height": 900},
|
||||
user_agent=(
|
||||
"Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) "
|
||||
"AppleWebKit/537.36 (KHTML, like Gecko) Chrome/126.0 Safari/537.36"
|
||||
),
|
||||
)
|
||||
context.set_default_timeout(15_000)
|
||||
return context, context.new_page()
|
||||
|
||||
|
||||
def run(args: argparse.Namespace) -> int:
|
||||
base_dir = Path(__file__).resolve().parent
|
||||
db_path = Path(args.db).expanduser().resolve()
|
||||
db_path.parent.mkdir(parents=True, exist_ok=True)
|
||||
with closing(sqlite3.connect(db_path)) as conn:
|
||||
init_db(conn, base_dir / "schema.sql")
|
||||
started_at = now_iso()
|
||||
run_id = conn.execute(
|
||||
"INSERT INTO crawl_runs(started_at) VALUES (?)", (started_at,)
|
||||
).lastrowid
|
||||
conn.commit()
|
||||
discovered_count = inserted_count = updated_count = failed_count = 0
|
||||
try:
|
||||
with sync_playwright() as playwright:
|
||||
context, page = new_context(playwright, args.headed is False)
|
||||
try:
|
||||
urls = discover_urls(page, args.list_url, args.scrolls, args.pause_ms)
|
||||
discovered_count = len(urls)
|
||||
inserted_count = insert_discoveries(conn, urls)
|
||||
queue = pending_urls(conn, urls, args.max_retries)
|
||||
cutoff = datetime.now() - timedelta(days=args.lookback_days)
|
||||
for index, url in enumerate(queue, start=1):
|
||||
path_date = parse_date(url.replace("/", "-"))
|
||||
if path_date and datetime.fromisoformat(path_date) < cutoff:
|
||||
continue
|
||||
try:
|
||||
article = extract_article(page, url)
|
||||
save_success(conn, article)
|
||||
updated_count += 1
|
||||
print(f"[{index}/{len(queue)}] OK {article['title']}")
|
||||
except Exception as exc:
|
||||
failed_count += 1
|
||||
save_failure(conn, url, exc)
|
||||
print(f"[{index}/{len(queue)}] FAIL {url}: {exc}", file=sys.stderr)
|
||||
time.sleep(args.article_pause)
|
||||
finally:
|
||||
context.close()
|
||||
status = "success" if failed_count == 0 else "partial"
|
||||
message = json.dumps({"db": str(db_path)}, ensure_ascii=False)
|
||||
except Exception as exc:
|
||||
status, message = "failed", str(exc)[:1000]
|
||||
print(f"crawl failed: {exc}", file=sys.stderr)
|
||||
conn.execute(
|
||||
"""
|
||||
UPDATE crawl_runs SET finished_at=?, status=?, discovered_count=?,
|
||||
inserted_count=?, updated_count=?, failed_count=?, message=?
|
||||
WHERE id=?
|
||||
""",
|
||||
(now_iso(), status, discovered_count, inserted_count, updated_count, failed_count, message, run_id),
|
||||
)
|
||||
conn.commit()
|
||||
print(json.dumps({
|
||||
"status": status, "discovered": discovered_count, "inserted": inserted_count,
|
||||
"updated": updated_count, "failed": failed_count, "database": str(db_path),
|
||||
}, ensure_ascii=False))
|
||||
return 0 if status in {"success", "partial"} else 1
|
||||
|
||||
|
||||
def build_parser() -> argparse.ArgumentParser:
|
||||
parser = argparse.ArgumentParser(description="红餐网文章每日增量采集入库")
|
||||
parser.add_argument("--db", default="hongcan.db", help="SQLite 数据库路径")
|
||||
parser.add_argument("--list-url", default=DEFAULT_LIST_URL)
|
||||
parser.add_argument("--lookback-days", type=int, default=7, help="回溯天数")
|
||||
parser.add_argument("--scrolls", type=int, default=12, help="列表页最大下拉次数")
|
||||
parser.add_argument("--pause-ms", type=int, default=1200, help="列表页下拉等待毫秒")
|
||||
parser.add_argument("--article-pause", type=float, default=0.8, help="文章间隔秒数")
|
||||
parser.add_argument("--max-retries", type=int, default=3)
|
||||
parser.add_argument("--headed", action="store_true", help="显示浏览器,便于排查验证码")
|
||||
return parser
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(run(build_parser().parse_args()))
|
||||
@@ -0,0 +1,4 @@
|
||||
playwright>=1.45,<2
|
||||
Flask>=3.0,<4
|
||||
python-dotenv>=1.0,<2
|
||||
|
||||
Executable
+4
@@ -0,0 +1,4 @@
|
||||
#!/usr/bin/env bash
|
||||
set -euo pipefail
|
||||
cd "$(dirname "$0")"
|
||||
exec python3 app.py
|
||||
+75
@@ -0,0 +1,75 @@
|
||||
PRAGMA journal_mode = WAL;
|
||||
PRAGMA foreign_keys = ON;
|
||||
|
||||
CREATE TABLE IF NOT EXISTS articles (
|
||||
id INTEGER PRIMARY KEY AUTOINCREMENT,
|
||||
canonical_url TEXT NOT NULL UNIQUE,
|
||||
source TEXT NOT NULL DEFAULT 'hongcan',
|
||||
title TEXT,
|
||||
author TEXT,
|
||||
published_at TEXT,
|
||||
summary TEXT,
|
||||
content TEXT,
|
||||
content_hash TEXT,
|
||||
crawl_status TEXT NOT NULL DEFAULT 'pending',
|
||||
retry_count INTEGER NOT NULL DEFAULT 0,
|
||||
last_error TEXT,
|
||||
discovered_at TEXT NOT NULL,
|
||||
crawled_at TEXT,
|
||||
updated_at TEXT NOT NULL
|
||||
);
|
||||
|
||||
CREATE INDEX IF NOT EXISTS idx_articles_published_at
|
||||
ON articles(published_at);
|
||||
CREATE INDEX IF NOT EXISTS idx_articles_crawl_status
|
||||
ON articles(crawl_status, retry_count);
|
||||
CREATE INDEX IF NOT EXISTS idx_articles_content_hash
|
||||
ON articles(content_hash);
|
||||
|
||||
CREATE TABLE IF NOT EXISTS crawl_runs (
|
||||
id INTEGER PRIMARY KEY AUTOINCREMENT,
|
||||
started_at TEXT NOT NULL,
|
||||
finished_at TEXT,
|
||||
status TEXT NOT NULL DEFAULT 'running',
|
||||
discovered_count INTEGER NOT NULL DEFAULT 0,
|
||||
inserted_count INTEGER NOT NULL DEFAULT 0,
|
||||
updated_count INTEGER NOT NULL DEFAULT 0,
|
||||
failed_count INTEGER NOT NULL DEFAULT 0,
|
||||
message TEXT
|
||||
);
|
||||
|
||||
CREATE TABLE IF NOT EXISTS ai_analyses (
|
||||
id INTEGER PRIMARY KEY AUTOINCREMENT,
|
||||
article_id INTEGER NOT NULL,
|
||||
model TEXT NOT NULL,
|
||||
analysis_type TEXT NOT NULL DEFAULT 'editorial',
|
||||
result_json TEXT NOT NULL,
|
||||
content_hash TEXT,
|
||||
created_at TEXT NOT NULL,
|
||||
updated_at TEXT NOT NULL,
|
||||
UNIQUE(article_id, model, analysis_type),
|
||||
FOREIGN KEY(article_id) REFERENCES articles(id) ON DELETE CASCADE
|
||||
);
|
||||
|
||||
CREATE INDEX IF NOT EXISTS idx_ai_analyses_article
|
||||
ON ai_analyses(article_id, updated_at);
|
||||
|
||||
CREATE TABLE IF NOT EXISTS industry_signals (
|
||||
id INTEGER PRIMARY KEY AUTOINCREMENT,
|
||||
signal_date TEXT NOT NULL,
|
||||
title TEXT NOT NULL,
|
||||
summary TEXT NOT NULL,
|
||||
trend TEXT NOT NULL DEFAULT 'emerging',
|
||||
confidence REAL NOT NULL DEFAULT 0.5,
|
||||
implications_json TEXT NOT NULL DEFAULT '[]',
|
||||
article_ids_json TEXT NOT NULL DEFAULT '[]',
|
||||
tags_json TEXT NOT NULL DEFAULT '[]',
|
||||
model TEXT NOT NULL,
|
||||
created_at TEXT NOT NULL,
|
||||
updated_at TEXT NOT NULL,
|
||||
UNIQUE(signal_date, title)
|
||||
);
|
||||
|
||||
CREATE INDEX IF NOT EXISTS idx_industry_signals_date
|
||||
ON industry_signals(signal_date DESC, confidence DESC);
|
||||
|
||||
@@ -0,0 +1,11 @@
|
||||
.article-card > p.ai-summary {
|
||||
color: #3f3e3a;
|
||||
}
|
||||
|
||||
.ai-summary b {
|
||||
display: inline-block;
|
||||
margin-right: 8px;
|
||||
color: var(--accent);
|
||||
font: 700 9px/1 ui-monospace, monospace;
|
||||
letter-spacing: .08em;
|
||||
}
|
||||
File diff suppressed because one or more lines are too long
@@ -0,0 +1,45 @@
|
||||
const button = document.querySelector('#analyzeBtn');
|
||||
const statusBox = document.querySelector('#analysisStatus');
|
||||
const content = document.querySelector('#analysisContent');
|
||||
|
||||
const escapeHtml = (value = '') => String(value).replace(/[&<>'"]/g, c => ({'&':'&','<':'<','>':'>',"'":''','"':'"'}[c]));
|
||||
|
||||
function renderList(items, ordered = false) {
|
||||
const tag = ordered ? 'ol' : 'ul';
|
||||
return `<${tag}>${(items || []).map(x => `<li>${escapeHtml(x)}</li>`).join('')}</${tag}>`;
|
||||
}
|
||||
|
||||
function renderAnalysis(a) {
|
||||
return `
|
||||
<section><h3>一句话摘要</h3><p>${escapeHtml(a.summary)}</p></section>
|
||||
<section><h3>核心要点</h3>${renderList(a.key_points)}</section>
|
||||
<section><h3>行业信号</h3><p>${escapeHtml(a.industry_signal)}</p></section>
|
||||
<section><h3>关键数据</h3>${renderList(a.numbers)}</section>
|
||||
<section><h3>延展选题</h3>${renderList(a.content_angles, true)}</section>
|
||||
<section><h3>风险提示</h3>${renderList(a.risk_notes)}</section>
|
||||
<section><h3>标签</h3><div class="tags">${(a.tags || []).map(x => `<span>${escapeHtml(x)}</span>`).join('')}</div></section>`;
|
||||
}
|
||||
|
||||
button?.addEventListener('click', async () => {
|
||||
button.disabled = true;
|
||||
statusBox.hidden = false;
|
||||
statusBox.className = 'analysis-status loading';
|
||||
statusBox.textContent = '千问正在阅读全文并提炼行业信号…';
|
||||
try {
|
||||
const response = await fetch(`/api/articles/${button.dataset.id}/analyze`, {
|
||||
method: 'POST', headers: {'Content-Type': 'application/json'}, body: JSON.stringify({force: button.textContent.includes('重新')})
|
||||
});
|
||||
const data = await response.json();
|
||||
if (!data.ok) throw new Error(data.error || '分析失败');
|
||||
content.innerHTML = renderAnalysis(data.analysis);
|
||||
statusBox.className = 'analysis-status success';
|
||||
statusBox.textContent = data.cached ? '已载入缓存分析' : '分析完成并已保存';
|
||||
button.textContent = '重新分析';
|
||||
} catch (error) {
|
||||
statusBox.className = 'analysis-status error';
|
||||
statusBox.textContent = error.message;
|
||||
} finally {
|
||||
button.disabled = false;
|
||||
}
|
||||
});
|
||||
|
||||
@@ -0,0 +1 @@
|
||||
.signals-block{padding:68px 0;border-bottom:1px solid var(--line)}.signals-title{display:flex;justify-content:space-between;align-items:end;margin-bottom:25px}.signals-title h2{margin:0;font-size:32px;letter-spacing:-.04em}.signals-title p,.signals-title time{margin:8px 0 0;color:var(--muted);font-size:12px}.signal-grid{display:grid;grid-template-columns:repeat(4,1fr);border-top:1px solid var(--ink)}.signal-card{min-height:300px;padding:22px;border-right:1px solid var(--line);border-bottom:1px solid var(--line)}.signal-card:nth-child(4n){border-right:0}.signal-card>div:first-child{display:flex;justify-content:space-between;align-items:center}.signal-card strong{font:11px ui-monospace,monospace;color:var(--muted)}.trend{display:inline-block;padding:5px 8px;border-radius:999px;background:var(--ink);color:white;font:700 9px ui-monospace,monospace;letter-spacing:.06em}.trend-加速{background:var(--accent)}.trend-新兴{background:#2b8a6f}.trend-稳定{background:#4a6fa5}.trend-下行{background:#888}.signal-card h3{margin:28px 0 14px;font-size:21px;line-height:1.3}.signal-card>p{color:var(--muted);font-size:13px;line-height:1.7}.signal-tags{display:flex;gap:5px;flex-wrap:wrap;margin-top:20px}.signal-tags span{font-size:10px;padding:4px 8px;border-radius:999px;background:rgba(216,73,47,.08);color:var(--accent);border:1px solid rgba(216,73,47,.15)}.signal-page{max-width:900px;padding:60px 0 100px}.signal-page header>p{color:var(--muted);font-size:12px}.signal-page h1{margin:22px 0;font-size:clamp(42px,6vw,76px);line-height:1.05;letter-spacing:-.055em}.signal-lead{font-size:20px;line-height:1.8;color:#454440}.signal-page section{margin-top:60px;border-top:1px solid var(--ink);padding-top:22px}.signal-page section h2{font-size:18px}.signal-page li{margin-bottom:14px;line-height:1.7}.evidence{display:grid;grid-template-columns:160px 1fr;gap:20px;padding:16px 0;border-bottom:1px solid var(--line)}.evidence time{color:var(--muted);font:11px ui-monospace,monospace}.evidence span{font-size:15px}.signal-page>.tags{margin-top:40px}@media(max-width:960px){.signal-grid{grid-template-columns:repeat(2,1fr)}.signal-card:nth-child(n){border-right:1px solid var(--line)}.signal-card:nth-child(even){border-right:0}}@media(max-width:640px){.signals-block{padding:44px 0}.signals-title{display:block;margin-bottom:18px}.signals-title h2{font-size:24px}.signals-title p,.signals-title time{font-size:11px;margin-top:6px}.signal-grid{grid-template-columns:1fr}.signal-card{min-height:auto;padding:20px 0;border-right:0}.signal-card h3{margin:20px 0 12px;font-size:19px;line-height:1.35}.signal-card>p{font-size:13px;line-height:1.65}.signal-tags{margin-top:16px}.signal-tags span{font-size:10px}.signal-page{padding:36px 0 50px}.signal-page header>p{font-size:11px}.signal-page h1{margin:18px 0;font-size:30px;line-height:1.12;letter-spacing:-.03em}.signal-lead{font-size:17px;line-height:1.7}.signal-page section{margin-top:40px;padding-top:18px}.signal-page section h2{font-size:16px}.signal-page li{margin-bottom:12px;line-height:1.65;font-size:15px}.evidence{grid-template-columns:1fr;gap:4px;padding:14px 0}.evidence time{font-size:10px}.evidence span{font-size:14px}.signal-page>.tags{margin-top:30px}}@media(max-width:480px){.signals-title h2{font-size:22px}.signal-card h3{font-size:18px}.signal-page h1{font-size:26px}.signal-lead{font-size:16px}}
|
||||
@@ -0,0 +1,36 @@
|
||||
{% extends "base.html" %}
|
||||
{% block title %}{{ article.title }} · 红餐观察库{% endblock %}
|
||||
{% block content %}
|
||||
<div class="reading-layout">
|
||||
<article class="story">
|
||||
<a class="back" href="/">← 返回文章库</a>
|
||||
<header>
|
||||
<div class="story-meta"><time>{{ article.published_at|date_cn }}</time><span>{{ article.author or '红餐网' }}</span></div>
|
||||
<h1>{{ article.title }}</h1>
|
||||
{% if article.summary %}<p class="dek">{{ article.summary }}</p>{% endif %}
|
||||
<a class="source-link" href="{{ article.canonical_url }}" target="_blank" rel="noopener">查看原文 ↗</a>
|
||||
</header>
|
||||
<div class="story-body">
|
||||
{% for paragraph in article.content.split('\n') if paragraph.strip() %}<p>{{ paragraph }}</p>{% endfor %}
|
||||
</div>
|
||||
</article>
|
||||
|
||||
<aside class="analysis-panel">
|
||||
<div class="analysis-head"><div><span>QWEN ANALYSIS</span><h2>编辑部研判</h2></div><button id="analyzeBtn" data-id="{{ article.id }}">{% if analysis %}重新分析{% else %}开始分析{% endif %}</button></div>
|
||||
<div id="analysisStatus" class="analysis-status" hidden></div>
|
||||
<div id="analysisContent">
|
||||
{% if analysis %}
|
||||
<section><h3>一句话摘要</h3><p>{{ analysis.summary }}</p></section>
|
||||
<section><h3>核心要点</h3><ul>{% for x in analysis.key_points %}<li>{{ x }}</li>{% endfor %}</ul></section>
|
||||
<section><h3>行业信号</h3><p>{{ analysis.industry_signal }}</p></section>
|
||||
<section><h3>延展选题</h3><ol>{% for x in analysis.content_angles %}<li>{{ x }}</li>{% endfor %}</ol></section>
|
||||
<section><h3>标签</h3><div class="tags">{% for x in analysis.tags %}<span>{{ x }}</span>{% endfor %}</div></section>
|
||||
{% else %}
|
||||
<div class="analysis-empty"><div class="scan-lines"></div><h3>等待 AI 阅读</h3><p>千问将提炼文章摘要、行业信号、关键数据和可延展选题。</p></div>
|
||||
{% endif %}
|
||||
</div>
|
||||
</aside>
|
||||
</div>
|
||||
{% endblock %}
|
||||
{% block scripts %}<script src="{{ url_for('static', filename='app.js') }}"></script>{% endblock %}
|
||||
|
||||
@@ -0,0 +1,24 @@
|
||||
<!doctype html>
|
||||
<html lang="zh-CN">
|
||||
<head>
|
||||
<meta charset="utf-8">
|
||||
<meta name="viewport" content="width=device-width,initial-scale=1">
|
||||
<meta name="color-scheme" content="light">
|
||||
<title>{% block title %}红餐观察库{% endblock %}</title>
|
||||
<link rel="stylesheet" href="{{ url_for('static', filename='app.css') }}">
|
||||
<link rel="stylesheet" href="{{ url_for('static', filename='ai-summary.css') }}">
|
||||
<link rel="stylesheet" href="{{ url_for('static', filename='signals.css') }}">
|
||||
</head>
|
||||
<body>
|
||||
<header class="site-header">
|
||||
<a class="brand" href="/" aria-label="返回文章库首页">
|
||||
<span class="brand-mark">HC</span>
|
||||
<span><b>红餐观察库</b><small>ARTICLE INTELLIGENCE</small></span>
|
||||
</a>
|
||||
<nav><a href="/">文章库</a><span class="system-dot"></span><span>本地运行</span></nav>
|
||||
</header>
|
||||
<main>{% block content %}{% endblock %}</main>
|
||||
<footer><span>红餐文章本地研究工具</span><span>SQLite · Playwright · Qwen</span></footer>
|
||||
{% block scripts %}{% endblock %}
|
||||
</body>
|
||||
</html>
|
||||
@@ -0,0 +1,61 @@
|
||||
{% extends "base.html" %}
|
||||
{% block content %}
|
||||
<section class="masthead">
|
||||
<div>
|
||||
<p class="kicker">餐饮行业内容情报</p>
|
||||
<h1>从文章堆里,<br>读出行业信号。</h1>
|
||||
</div>
|
||||
<form class="search" method="get" role="search">
|
||||
<label for="q">搜索标题、品牌或正文</label>
|
||||
<div><input id="q" name="q" value="{{ query }}" placeholder="例如:咖啡、海底捞、供应链"><button>搜索</button></div>
|
||||
<p>已收录 {{ stats.total }} 篇,{{ stats.analyzed or 0 }} 篇完成 AI 分析</p>
|
||||
</form>
|
||||
</section>
|
||||
|
||||
<section class="metrics" aria-label="文章库数据">
|
||||
<div><strong>{{ stats.success }}</strong><span>完整文章</span></div>
|
||||
<div><strong>{{ stats.analyzed or 0 }}</strong><span>AI 分析</span></div>
|
||||
<div><strong>{{ stats.latest|date_cn }}</strong><span>最近更新</span></div>
|
||||
</section>
|
||||
|
||||
{% if signals and not query and page == 1 %}
|
||||
<section class="signals-block">
|
||||
<div class="signals-title"><div><h2>今日行业信号</h2><p>基于多篇文章交叉提炼,不是单篇新闻复述</p></div><time>{{ signals[0].signal_date }}</time></div>
|
||||
<div class="signal-grid">
|
||||
{% for signal in signals %}
|
||||
<a class="signal-card" href="/signals/{{ signal.id }}">
|
||||
<div><span class="trend trend-{{ signal.trend }}">{{ signal.trend }}</span><strong>{{ (signal.confidence * 100)|round|int }}%</strong></div>
|
||||
<h3>{{ signal.title }}</h3><p>{{ signal.summary }}</p>
|
||||
<div class="signal-tags">{% for tag in signal.tags[:3] %}<span>{{ tag }}</span>{% endfor %}</div>
|
||||
</a>
|
||||
{% endfor %}
|
||||
</div>
|
||||
</section>
|
||||
{% endif %}
|
||||
|
||||
<section class="library-head">
|
||||
<div><h2>{% if query %}“{{ query }}”的结果{% else %}最新入库{% endif %}</h2><p>共 {{ total }} 篇</p></div>
|
||||
{% if query %}<a href="/">清除搜索</a>{% endif %}
|
||||
</section>
|
||||
|
||||
<section class="article-grid">
|
||||
{% for item in articles %}
|
||||
<article class="article-card">
|
||||
<div class="card-meta"><time>{{ item.published_at|date_cn }}</time>{% if item.analyzed %}<span>AI 已分析</span>{% endif %}</div>
|
||||
<h3><a href="/article/{{ item.id }}">{{ item.title }}</a></h3>
|
||||
{% if item.ai_summary %}<p class="ai-summary"><b>AI 摘要</b>{{ item.ai_summary }}</p>{% else %}<p>{{ item.summary or item.excerpt }}</p>{% endif %}
|
||||
<a class="read-link" href="/article/{{ item.id }}">阅读全文 <span>↗</span></a>
|
||||
</article>
|
||||
{% else %}
|
||||
<div class="empty"><h3>没有找到相关文章</h3><p>换一个品牌名、品类或行业关键词试试。</p></div>
|
||||
{% endfor %}
|
||||
</section>
|
||||
|
||||
{% if pages > 1 %}
|
||||
<nav class="pagination" aria-label="分页">
|
||||
{% if page > 1 %}<a href="?q={{ query|urlencode }}&page={{ page-1 }}">上一页</a>{% endif %}
|
||||
<span>{{ page }} / {{ pages }}</span>
|
||||
{% if page < pages %}<a href="?q={{ query|urlencode }}&page={{ page+1 }}">下一页</a>{% endif %}
|
||||
</nav>
|
||||
{% endif %}
|
||||
{% endblock %}
|
||||
@@ -0,0 +1,11 @@
|
||||
{% extends "base.html" %}
|
||||
{% block title %}{{ signal.title }} · 行业信号{% endblock %}
|
||||
{% block content %}
|
||||
<article class="signal-page">
|
||||
<a class="back" href="/">← 返回行业信号</a>
|
||||
<header><span class="trend trend-{{ signal.trend }}">{{ signal.trend }}</span><p>{{ signal.signal_date }} · 置信度 {{ (signal.confidence*100)|round|int }}%</p><h1>{{ signal.title }}</h1><div class="signal-lead">{{ signal.summary }}</div></header>
|
||||
<section><h2>经营启示</h2><ol>{% for item in implications %}<li>{{ item }}</li>{% endfor %}</ol></section>
|
||||
<section><h2>证据文章</h2>{% for item in articles %}<a class="evidence" href="/article/{{ item.id }}"><time>{{ item.published_at|date_cn }}</time><span>{{ item.title }}</span></a>{% endfor %}</section>
|
||||
<div class="tags">{% for tag in tags %}<span>{{ tag }}</span>{% endfor %}</div>
|
||||
</article>
|
||||
{% endblock %}
|
||||
Reference in New Issue
Block a user