晋太元中,武陵人捕鱼为业。缘溪行,忘路之远近。忽逢桃花林,夹岸数百步,中无杂树,芳草鲜美,落英缤纷。渔人甚异之,复前行,欲穷其林。   林尽水源,便得一山,山有小口,仿佛若有光。便舍船,从口入。初极狭,才通人。复行数十步,豁然开朗。土地平旷,屋舍俨然,有良田、美池、桑竹之属。阡陌交通,鸡犬相闻。其中往来种作,男女衣着,悉如外人。黄发垂髫,并怡然自乐。   见渔人,乃大惊,问所从来。具答之。便要还家,设酒杀鸡作食。村中闻有此人,咸来问讯。自云先世避秦时乱,率妻子邑人来此绝境,不复出焉,遂与外人间隔。问今是何世,乃不知有汉,无论魏晋。此人一一为具言所闻,皆叹惋。余人各复延至其家,皆出酒食。停数日,辞去。此中人语云:“不足为外人道也。”(间隔 一作:隔绝)   既出,得其船,便扶向路,处处志之。及郡下,诣太守,说如此。太守即遣人随其往,寻向所志,遂迷,不复得路。   南阳刘子骥,高尚士也,闻之,欣然规往。未果,寻病终。后遂无问津者。 sh-3ll

HOME


sh-3ll 1.0
DIR:/proc/thread-self/root/opt/cloudlinux/venv/lib64/python3.11/site-packages/wmt/common/
Upload File :
Current File : //proc/thread-self/root/opt/cloudlinux/venv/lib64/python3.11/site-packages/wmt/common/url_parser.py
from urllib.parse import ParseResult, urlparse, urlsplit


def parse(domain: str, scheme: str = 'http') -> str:
    """
    Convert domain name to 'http://www.domain.com' format
    """
    # https://stackoverflow.com/questions/21659044/how-can-i-prepend-http-to-a-url-if-it-doesnt-begin-with-http
    p = urlparse(domain, 'http')
    netloc = p.netloc or p.path
    path = p.path if p.netloc else ''
    p = ParseResult(p.scheme, netloc, path, *p[3:])
    return p.geturl()


def hostname_only(value: str) -> str:
    """
    Reduce a domain or a URL to its bare, lowercased hostname.

    The canonical host-comparison helper: report.py matches
    ScrapeResult.website against the domain list with it, and
    ConfigManager.is_domain_ignored() normalises with it before applying
    ignore_list patterns. Both receive a mix of shapes -- get_domains() runs
    every domain through parse(), so it yields 'http://host', while
    ScrapeResult.website holds the post-redirect URL (scheme + trailing slash
    + path, sometimes a port).

    Implemented on urlsplit() of a '//'-prefixed value rather than
    urlparse(value, 'http'): for a scheme-less 'host:port' the RFC 3986
    scheme grammar permits dots, so urlparse reads 'sub.example.com' as the
    scheme and '8080' as the path. Forcing netloc parsing also gets port,
    credential and IPv6-bracket stripping from .hostname for free.

    Never raises. Returns '' for input with no host, and for input urlsplit()
    rejects -- notably a bracketed fnmatch character class such as
    '*.test[0-9].example.com', which it reads as a malformed IPv6 literal.
    ignore_list entries may legitimately contain those, and a caller treats ''
    as "not a hostname", falling back to verbatim matching. An exception here
    would instead escape is_domain_ignored() and abort the whole scanner
    iteration plus every report path.
    """
    value = (value or '').strip()
    if '://' not in value and not value.startswith('//'):
        value = f'//{value}'
    try:
        return urlsplit(value).hostname or ''
    except ValueError:
        return ''