> ## Documentation Index
> Fetch the complete documentation index at: https://leetcode-py.wisl.dev/llms.txt
> Use this file to discover all available pages before exploring further.

> ## Agent Instructions
> leetcode-py is a Python LeetCode practice environment generator with one CLI: lcpy. It is not a service or platform.
> Each problem is a directory under leetcode/ with README.md, solution.py, test_solution.py, helpers.py, and playground.ipynb. lcpy gen creates them from JSON templates bundled with the package.
> Examples are backed by tests; copy them verbatim.

# Web Crawler Python Solution with Tests

> Tested Python solution for LeetCode 1236 with 18 pytest cases. Generate a practice environment with lcpy.

LeetCode 1236, [Medium](/catalog/medium). Topics: [Depth-First Search](/catalog/topics/depth-first-search), [Breadth-First Search](/catalog/topics/breadth-first-search), [String](/catalog/topics/string), [Interactive](/catalog/topics/interactive). [View on LeetCode](https://leetcode.com/problems/web-crawler/description/).

Generate this problem as a practice environment: tested reference solution, 18 [parametrized pytest cases](/practice/testing), and a playground notebook:

```bash theme={"theme":{"light":"github-light","dark":"github-dark"}}
lcpy gen -n 1236   # by problem number
lcpy gen -s web_crawler   # by problem name
```

## Problem

Given a url `startUrl` and an interface `HtmlParser`, implement a web crawler to crawl all links that are under the same hostname as `startUrl`.

Return all urls obtained by your web crawler in **any** order.

Your crawler should:

* Start from the page: `startUrl`
* Call `HtmlParser.getUrls(url)` to get all urls from a webpage of given url.
* Do not crawl the same link twice.
* Explore only the links that are under the same hostname as `startUrl`.

![Hostname](https://fastly.jsdelivr.net/gh/doocs/leetcode@main/solution/1200-1299/1236.Web%20Crawler/images/urlhostname.png)

As shown in the example url above, the hostname is `example.org`. For simplicity sake, you may assume all urls use http protocol without any port specified. For example, the urls `http://leetcode.com/problems` and `http://leetcode.com/contest` are under the same hostname, while urls `http://example.org/test` and `http://example.com/abc` are not under the same hostname.

The `HtmlParser` interface is defined as such:

```
interface HtmlParser {
  // Return a list of all urls from a webpage of given url.
  public List<String> getUrls(String url);
}
```

Below are two examples explaining the functionality of the problem, for custom testing purposes you'll have three variables `urls`, `edges` and `startUrl`. Notice that you will only have access to `startUrl` in your code, while `urls` and `edges` are not directly accessible to you in code.

Note: Consider the same URL with the trailing slash `/` as a different URL. For example, `http://news.yahoo.com`, and `http://news.yahoo.com/` are different urls.

### Examples

![Example 1](https://fastly.jsdelivr.net/gh/doocs/leetcode@main/solution/1200-1299/1236.Web%20Crawler/images/sample_2_1497.png)

```
Input:
urls = [
  "http://news.yahoo.com",
  "http://news.yahoo.com/news",
  "http://news.yahoo.com/news/topics/",
  "http://news.google.com",
  "http://news.yahoo.com/us"
]
edges = [[2,0],[2,1],[3,2],[3,1],[0,4]]
startUrl = "http://news.yahoo.com/news/topics/"
Output:
[
  "http://news.yahoo.com",
  "http://news.yahoo.com/news",
  "http://news.yahoo.com/news/topics/",
  "http://news.yahoo.com/us"
]
```

![Example 2](https://fastly.jsdelivr.net/gh/doocs/leetcode@main/solution/1200-1299/1236.Web%20Crawler/images/sample_3_1497.png)

```
Input:
urls = [
  "http://news.yahoo.com",
  "http://news.yahoo.com/news",
  "http://news.yahoo.com/news/topics/",
  "http://news.google.com"
]
edges = [[0,2],[2,1],[3,2],[3,1],[3,0]]
startUrl = "http://news.google.com"
Output: ["http://news.google.com"]
Explanation: The startUrl links to all other pages that do not share the same hostname.
```

### Constraints

* `1 <= urls.length <= 1000`
* `1 <= urls[i].length <= 300`
* `startUrl` is one of the `urls`.
* Hostname label must be from 1 to 63 characters long, including the dots, may contain only the ASCII letters from 'a' to 'z', digits from '0' to '9' and the hyphen-minus character ('-').
* The hostname may not start or end with the hyphen-minus character ('-').
* You may assume there're no duplicates in url library.

## Solution

Reference implementation from [solution.py on GitHub](https://github.com/wislertt/leetcode-py/blob/main/leetcode/web_crawler/solution.py), full suite in [test\_solution.py](https://github.com/wislertt/leetcode-py/blob/main/leetcode/web_crawler/test_solution.py):

```python theme={"theme":{"light":"github-light","dark":"github-dark"}}
class HtmlParser:
    # Test-harness API: backs the getUrls interface with the url library
    def __init__(self, urls: list[str], edges: list[list[int]]) -> None:
        self.adj: dict[str, list[str]] = {u: [] for u in urls}
        for a, b in edges:
            self.adj[urls[a]].append(urls[b])

    def get_urls(self, url: str) -> list[str]:
        return self.adj[url]


class Solution:
    # Time: O(V + E) pages and links visited once
    # Space: O(V) visited set
    def crawl(self, start_url: str, html_parser: HtmlParser) -> list[str]:
        host = start_url[7:].split("/")[0]
        visited = {start_url}
        stack = [start_url]
        while stack:
            url = stack.pop()
            for link in html_parser.get_urls(url):
                if link in visited or link[7:].split("/")[0] != host:
                    continue
                visited.add(link)
                stack.append(link)
        return list(visited)
```

## Complexity

| Time | Space |
| - | - |
| O(V + E) pages and links visited once | O(V) visited set |

## Tags

[NeetCode All](/catalog/neetcode).


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.