Logo of «2Captcha»To home page
Captcha bypass tutorials

Was this helpful?

How to scrape SHEIN: A Practical guide

Gregory Fisher
Gregory Fisher

Technical engineer

SHEIN is a major e-commerce platform in the fashion industry. Scraping its search pages, categories, and product cards is more complex than simply loading HTML and searching for elements by CSS selectors.

The main complexity lies in two factors:

  1. Product data is dynamically embedded in the JavaScript state of the page (window.gbRawData for listings) or in schema.org JSON-LD (for product cards), rather than being rendered in plain HTML.
  2. SHEIN uses a proprietary risk scoring system that can redirect automated sessions to /risk/challenge (interactive verification) or /risk/action/limit (hard rate limiting).

In this guide, we'll examine how the open-source project 2scraper/shein-scraper works. We'll focus on the Playwright engine, as it's the recommended and most stable one in this repository.


What We'll Build

A scraper capable of working in three modes:

  • Searching for products by keyword.
  • Scraping a specific SHEIN category.
  • Scraping an individual product page.

Extracted data includes: sku, title, brand, price, original_price, currency, discount_pct, rating, review_count, in_stock, is_clearance, quickship, and product_url.

Important: The current implementation is focused on us.shein.com. You should not automatically assume that other regional domains use identical data structures or endpoints.


Requirements

  • Python: 3.9 or newer.
  • Browser engine: Playwright (recommended). Dependencies for Selenium and Puppeteer are separated in the project, as their packages may conflict.
  • External services (optional):
    • 2Captcha API key (TWOCAPTCHA_KEY) if automatic solving of visual challenge tasks is required.
    • Scraping Browser CDP endpoint (SHEIN_CDP_ENDPOINT) for using persistent browser profiles.

Project Setup

  1. Clone the repository and set up the environment:

    bash Copy
    git clone https://github.com/2scraper/shein-scraper.git
    cd shein-scraper
    python3 -m venv .venv
    source .venv/bin/activate  # For Windows: .venv\Scripts\activate
  2. Install dependencies and the browser:

    bash Copy
    pip install -r requirements-playwright.txt
    playwright install chromium
  3. Configure environment variables:

    bash Copy
    cp .env.example .env

    Edit .env to add TWOCAPTCHA_KEY, SHEIN_PROXY, or SHEIN_CDP_ENDPOINT as needed. Storing credentials in .env is preferable to passing them via command-line arguments.


Quick Start

The simplest way to run the scraper is via CLI:

bash Copy
# Search by keyword
python3 playwright_scraper.py --query "summer dress" --max-results 20

# Scrape a category (path taken from the actual SHEIN URL, e.g., Women-Jeans-c-1934.html)
python3 playwright_scraper.py --category "Women-Jeans-c-1934.html" --format csv --out jeans.csv

# Scrape an individual product card
python3 playwright_scraper.py --url "https://us.shein.com/dsbayvkj-p-33704388.html"

By default, results are saved to shein_results.json. A metadata file shein_results.json.meta.json is created alongside it, which is critical for automation: it allows you to distinguish between "there are truly no products" and "the scraper was blocked."


How the Scraper Works

The repository architecture separates the browser engine from the parsing logic. The files playwright_scraper.py, selenium_scraper.py, and puppeteer_scraper.py use the common execution logic from page_flow.py.

Simplified execution flow:

  1. URL formation (based on the actual SHEIN structure, not invented internal APIs).
  2. Launching the browser and navigating to the page.
  3. Instant URL check: if the address contains /risk/challenge or /risk/action/limit, the session is marked as blocked.
  4. Waiting and scrolling the page (for listings).
  5. Extracting window.gbRawData (preferably via page.evaluate, with a fallback to HTML parsing).
  6. Parsing, deduplication, and saving results.

Why use a browser instead of requests?
Because the main data source is a JavaScript object. The browser adapter can retrieve it directly (() => window.gbRawData), which is much more reliable than trying to parse complex, deeply nested JSON from raw HTML using regular expressions.


Data Extraction

1. Search and Categories (Primary Path)

Data is extracted from window.gbRawData. The scraper reads the raw product records, not the text from the page:

python Copy
goods_id = raw.get("goods_id")
price = _money(raw.get("salePrice"))
original_price = _money(raw.get("retailPrice"))

The scraper also extracts the stated total number of results (sum or result_count). This allows the program to understand whether the collection was intentionally stopped by the --max-results parameter or if the scraper unexpectedly received fewer products than the site reported.

2. Individual Product Page

A different path is used for product cards. The scraper looks for <script type="application/ld+json"> and parses schema.org ProductGroup or Product.

Important: The repository does not use JSON-LD for search pages, as it only contains navigation data (breadcrumb) there, not the product list.

3. DOM Parsing (Fallback)

The code contains fallback CSS selectors (e.g., .product-card, [data-goods-id]). However, in the source code, they are explicitly marked with the comment # TODO: verify live. This means they are a best-effort solution in case window.gbRawData disappears and should not be considered a confirmed primary extraction method.


Working with SHEIN's Anti-Bot Protection

The repository clearly separates two types of protective mechanisms, and the scraper reacts to them differently.

/risk/challenge (Interactive Verification)

This is a session verification gateway. The implementation in the repository recognizes SHEIN-specific visual tasks rather than trying to pass them off as standard reCAPTCHA:

  1. 3×3 Grid CAPTCHA: The scraper takes a screenshot, encodes it in Base64, and sends it as a GridTask. In response, it receives cell numbers (e.g., [1, 4, 7]), converts them to coordinates, and clicks on them.
  2. Icon Sequence CAPTCHA: Sent as a CoordinatesTask. The response contains an array of coordinates {"x": 120, "y": 83}, which the scraper scales to the CSS dimensions of the browser window (important for high-DPI displays) and emulates clicks.

After clicking, the scraper checks if the challenge page has disappeared. If not, the task can be updated and solved again (up to the --risk-challenge-rounds limit).

/risk/action/limit (Rate Limit)

This is not a CAPTCHA. This is a hard rate limit. Sending such a page to a CAPTCHA solving service is pointless. The scraper recognizes this path, logs it, and can wait for a specified time (--rate-limit-cooldown 300) before finishing or retrying.


Working with CAPTCHA (2Captcha Integration)

If the TWOCAPTCHA_KEY is enabled, the scraper automatically delegates the solving of visual tasks. The request lifecycle for GridTask looks like this:

python Copy
import base64
import requests

API_KEY = "YOUR_2CAPTCHA_KEY"

with open("challenge_grid.png", "rb") as f:
    body = base64.b64encode(f.read()).decode()

# 1. Creating the task
response = requests.post(
    "https://api.2captcha.com/createTask",
    json={
        "clientKey": API_KEY,
        "task": {
            "type": "GridTask",
            "body": body,
            "rows": 3,
            "columns": 3,
            "comment": "Select the images matching the instruction"
        }
    }
)
task_id = response.json()["taskId"]

# 2. Waiting for the result
while True:
    result = requests.post(
        "https://api.2captcha.com/getTaskResult",
        json={"clientKey": API_KEY, "taskId": task_id}
    ).json()
    
    if result.get("status") == "ready":
        clicks = result["solution"]["click"] # For example: [1, 4, 7]
        break
    time.sleep(5)
    
# 3. Converting clicks to coordinates and emulating clicks in the browser

Important: Getting the correct answer from the CAPTCHA solving API does not guarantee that SHEIN will accept the session. If the risk scoring remains high, the challenge may appear again. The --max-solves parameter protects against infinite billing.


Using Proxies

Proxy support is built into the repository and is not just a theoretical recommendation. The format http://user:pass@host:port or the compact host:port:login:password is supported.

bash Copy
python3 playwright_scraper.py --query "jeans" --proxy-file proxies.txt

The internal proxy_pool.py uses round-robin selection and tracks recurring errors, temporarily excluding non-working proxies. In Playwright, credentials are passed through the standard configuration, which is more reliable than injecting them into Chromium launch arguments.


Using 2Captcha Browser API (CDP)

The Playwright implementation supports connecting to a remote browser via Chrome DevTools Protocol (CDP) instead of launching a local Chromium:

bash Copy
python3 playwright_scraper.py --query "summer dress" --cdp-endpoint "$SHEIN_CDP_ENDPOINT"

When this is preferable:

  • Local sessions constantly hit /risk/challenge.
  • You need to preserve cookies and browser state between runs (persistent profile).
  • You need to separate browser infrastructure from scraper logic.

The remote browser itself does not give a 100% guarantee of access, but it radically reduces the frequency of risk gateway appearances compared to creating a new "clean" profile on each run.


Scaling the Scraper

The repository provides basic mechanisms (retry, jitter, proxy rotation), but orchestration is required for production loads:

  1. Limit parallelism: Don't launch hundreds of browsers simultaneously. Start with a small pool and monitor the frequency of challenge appearances. The repository does not contain confirmed "safe" request limits for SHEIN.
  2. Preserve sessions: For periodic monitoring, reusing a single CDP profile is more efficient than constantly generating new ones.
  3. Add jitter: The --delay-jitter and --scroll-delay parameters prevent synchronous request spikes from multiple workers.
  4. Check completeness: Compare the number of collected products with the result_count field from window.gbRawData. If the scraper collected 20 products and the site reports 1500, the run should be considered partial (especially considering that loading additional data during scrolling is marked in the repo as requiring additional live verification).

Troubleshooting

Problem Possible Cause Solution
Redirect to /risk/challenge SHEIN requested session verification Use the built-in challenge handler. If the profile is constantly rejected, change the CDP endpoint or proxy.
Opening /risk/action/limit Rate limit triggered Do not send this to a CAPTCHA solver. Reduce request frequency, use --rate-limit-cooldown, or retry later.
Returns 0 products The output is empty, or data didn't load Check the .meta.json file. Use --dump-html to see if the site returned a consent page or error.
Only the first batch of products is collected (~20 items) Scroll-driven loading of large catalogs is not fully confirmed in the current version Consider large samples as potentially partial. For full category scraping, use pagination via real subcategory URLs.
CAPTCHA solved, but challenge remains SHEIN rejected the session despite the correct solution to the visual task Get a new task (within --max-solves) or change the profile/proxy.
DOM fallback stopped working Selectors are marked as TODO: verify live Don't rely on them. The main source is window.gbRawData.

Example Output

The price_source field in the output indicates the origin of the data, which is critical for auditing:

json Copy
{
  "sku": "458057728",
  "source": "shein.com",
  "category": "Women Mini Dresses",
  "title": "Aloruh Women's Solid Color Sleeveless Mini Dress",
  "brand": "Aloruh",
  "price": 13.03,
  "currency": "USD",
  "price_source": "embedded_json",
  "product_url": "https://us.shein.com/Aloruh-Women-s-Solid-Color-Sleeveless-Mini-Dress-p-458057728.html",
  "image_url": "https://img.ltwebstatic.com/v4/j/pi/.../thumbnail_405x552.jpg",
  "scraped_at": "2026-09-30T12:56:49Z",
  "original_price": 20.89,
  "discount_pct": 38.0,
  "rating": 4.62,
  "review_count": 1001,
  "in_stock": true,
  "is_clearance": false,
  "quickship": false
}

Note: For individual product pages, price_source can take the values json_ld or json_ld_min_variant (if the price is taken as the minimum among variants).


Complete Compact Code Example

Below is a minimal working example demonstrating the repository's key technique: retrieving window.gbRawData via Playwright with basic error handling.

python Copy
import asyncio
from urllib.parse import quote
from playwright.async_api import async_playwright

BASE_URL = "https://us.shein.com"

def search_url(query: str) -> str:
    return f"{BASE_URL}/pdsearch/{quote(query)}/"

def parse_products(data: dict, limit: int = 20) -> list[dict]:
    results = data.get("results") or {}
    info = results.get("bffProductsInfo") or {}
    raw_products = info.get("products") or []

    products = []
    for raw in raw_products[:limit]:
        goods_id = raw.get("goods_id")
        if not goods_id:
            continue
            
        products.append({
            "sku": str(goods_id),
            "title": raw.get("goods_name"),
            "price": raw.get("salePrice", {}).get("amount"),
            "product_url": f"{BASE_URL}/{raw.get('goods_url_name') or 'product'}-p-{goods_id}.html",
        })
    return products

async def main():
    async with async_playwright() as p:
        # In production, you can add proxy={"server": "..."} here
        browser = await p.chromium.launch(headless=True)
        context = await browser.new_context()
        page = await context.new_page()

        url = search_url("summer dress")
        print(f"Navigating to URL: {url}")
        
        await page.goto(url, wait_until="domcontentloaded", timeout=60_000)

        # Check for risk gateways
        if "/risk/challenge" in page.url or "/risk/action/limit" in page.url:
            print(f"Session blocked or limited: {page.url}")
            await browser.close()
            return

        # Preferred extraction method
        data = await page.evaluate("() => window.gbRawData || null")

        if not data:
            print("window.gbRawData not found. Possible anti-bot page or frontend change.")
            await browser.close()
            return

        products = parse_products(data, limit=5)
        print(f"Successfully extracted {len(products)} products:")
        for p in products:
            print(f" - {p['sku']}: {p['title']} ({p['price']} USD)")

        await browser.close()

if __name__ == "__main__":
    asyncio.run(main())

This example intentionally does not include challenge handling, proxy rotation, and scrolling to demonstrate exactly the core of data extraction.


Testing and Limitations

The repository contains offline smoke tests (python3 smoke_test.py) that verify the deterministic parsing logic and CAPTCHA task construction.

However, it's important to clearly separate offline tests and the live behavior of the site. According to the project documentation, at the time of the last checks, the following were confirmed:

  • Playwright working with window.gbRawData.
  • Using JSON-LD for product cards.
  • Recognition of the /risk/challenge and /risk/action/limit paths.

Areas requiring additional live verification (marked in the code as TODO: verify live):

  • Loading additional results when scrolling beyond the first batch.
  • Functionality of DOM fallback selectors.
  • End-to-end scenarios for Selenium and Puppeteer.

You should not turn these unconfirmed capabilities into guaranteed statements when building a production pipeline.


Conclusion

Reliable SHEIN scraping is not about searching for "magic" CSS selectors, but understanding where the frontend gets structured data from. Using window.gbRawData for listings and JSON-LD for product cards makes the scraper resistant to cosmetic changes in the layout.

Anti-bot states should be considered part of the standard workflow: /risk/challenge is handled via visual tasks (Grid/Coordinates), while /risk/action/limit requires strategic waiting, not attempts to solve CAPTCHA.

For one-off tasks, local Playwright is sufficient. For regular price monitoring, it is highly recommended to use persistent browser sessions (via CDP), proxies, and the built-in diff_runs.py tool for comparing results, which allows you to avoid false conclusions about products "disappearing" due to temporary scraper blocking.