晋太元中,武陵人捕鱼为业。缘溪行,忘路之远近。忽逢桃花林,夹岸数百步,中无杂树,芳草鲜美,落英缤纷。渔人甚异之,复前行,欲穷其林。 林尽水源,便得一山,山有小口,仿佛若有光。便舍船,从口入。初极狭,才通人。复行数十步,豁然开朗。土地平旷,屋舍俨然,有良田、美池、桑竹之属。阡陌交通,鸡犬相闻。其中往来种作,男女衣着,悉如外人。黄发垂髫,并怡然自乐。 见渔人,乃大惊,问所从来。具答之。便要还家,设酒杀鸡作食。村中闻有此人,咸来问讯。自云先世避秦时乱,率妻子邑人来此绝境,不复出焉,遂与外人间隔。问今是何世,乃不知有汉,无论魏晋。此人一一为具言所闻,皆叹惋。余人各复延至其家,皆出酒食。停数日,辞去。此中人语云:“不足为外人道也。”(间隔 一作:隔绝) 既出,得其船,便扶向路,处处志之。及郡下,诣太守,说如此。太守即遣人随其往,寻向所志,遂迷,不复得路。 南阳刘子骥,高尚士也,闻之,欣然规往。未果,寻病终。后遂无问津者。
| DIR:/proc/thread-self/root/opt/cloudlinux/venv/lib/python3.11/site-packages/wmt/common/ |
| Current File : //proc/thread-self/root/opt/cloudlinux/venv/lib/python3.11/site-packages/wmt/common/url_parser.py |
from urllib.parse import ParseResult, urlparse, urlsplit
def parse(domain: str, scheme: str = 'http') -> str:
"""
Convert domain name to 'http://www.domain.com' format
"""
# https://stackoverflow.com/questions/21659044/how-can-i-prepend-http-to-a-url-if-it-doesnt-begin-with-http
p = urlparse(domain, 'http')
netloc = p.netloc or p.path
path = p.path if p.netloc else ''
p = ParseResult(p.scheme, netloc, path, *p[3:])
return p.geturl()
def hostname_only(value: str) -> str:
"""
Reduce a domain or a URL to its bare, lowercased hostname.
The canonical host-comparison helper: report.py matches
ScrapeResult.website against the domain list with it, and
ConfigManager.is_domain_ignored() normalises with it before applying
ignore_list patterns. Both receive a mix of shapes -- get_domains() runs
every domain through parse(), so it yields 'http://host', while
ScrapeResult.website holds the post-redirect URL (scheme + trailing slash
+ path, sometimes a port).
Implemented on urlsplit() of a '//'-prefixed value rather than
urlparse(value, 'http'): for a scheme-less 'host:port' the RFC 3986
scheme grammar permits dots, so urlparse reads 'sub.example.com' as the
scheme and '8080' as the path. Forcing netloc parsing also gets port,
credential and IPv6-bracket stripping from .hostname for free.
Never raises. Returns '' for input with no host, and for input urlsplit()
rejects -- notably a bracketed fnmatch character class such as
'*.test[0-9].example.com', which it reads as a malformed IPv6 literal.
ignore_list entries may legitimately contain those, and a caller treats ''
as "not a hostname", falling back to verbatim matching. An exception here
would instead escape is_domain_ignored() and abort the whole scanner
iteration plus every report path.
"""
value = (value or '').strip()
if '://' not in value and not value.startswith('//'):
value = f'//{value}'
try:
return urlsplit(value).hostname or ''
except ValueError:
return ''
|