A production-ready, feature-rich web scraping framework that bypasses common anti-bot measures using advanced evasion techniques.
- Automatic rotation with each session
- 5+ default modern user agents
- Support for custom user agent lists
- Round-robin proxy switching
- Support for HTTP, HTTPS, SOCKS5
- Authenticated proxy support
- Tor Browser integration
- Random delays (configurable ranges)
- Natural typing with realistic typos
- Human-like mouse movements
- Realistic scrolling patterns
- Reading delays based on content length
- Automatic cookie cleaning
- Fresh session with each proxy rotation
- Configurable requests-per-session limit
- Detects invisible elements
- Pattern-based detection (class/id/name)
- Skips trap links automatically
safe_click()andget_safe_links()methods
- Canvas fingerprint randomization
- Audio context spoofing
- WebGL parameter masking
- Randomized viewport sizes
- Timezone randomization
- Removes
navigator.webdriver - Mocks Chrome runtime
- Spoofs plugins & languages
- Removes automation indicators
- Easy Tor integration (
use_tor=True) - SOCKS5 proxy configuration
- Works with Tor Browser bundle
pip install playwright
playwright install chromium firefox
# User agent
user_agents=[], # Custom user agents (optional)
rotate_user_agent=True, # Rotate UA each session
# Human behavior
typing_speed_range=(50, 150), # Milliseconds per keystroke
action_delay_range=(0.5, 2.0), # Seconds between actions
reading_delay_range=(1.0, 3.0), # Reading time in seconds
enable_human_behavior=True, # Enable human simulation
# Session management
requests_per_session=10, # Requests before rotation
auto_rotate_session=True, # Auto-rotate sessions
clear_cookies=True, # Clear cookies on rotation
# Detection evasion
detect_honeypots=True, # Enable honeypot detection
stealth_mode=True, # Enable stealth techniques
randomize_viewport=True, # Random screen sizes
randomize_timezone=True, # Random timezones
# Browser
headless=False, # Headless mode
browser_type="chromium" # chromium, firefox, webkit
The scraper includes 6 ready-to-use project templates:
Track product prices across multiple stores
project_1_ecommerce_price_monitor()
Collect articles from news sites with honeypot avoidance
project_2_news_aggregator()
Extract property listings with details
project_3_real_estate_scraper()
Auto-fill and submit forms with human behavior
project_4_form_automation()
Collect job postings from multiple sources
project_5_job_board_scraper()
Track trends (respecting ToS)
project_6_social_media_monitor()
def your_custom_scraper(): # Step 1: Configure config = ScraperConfig( proxies=[ProxyServer("http://proxy:8080")], enable_human_behavior=True, detect_honeypots=True, headless=False )
# Step 2: Create scraper
scraper = AntiBlockingScraper(config)
# Step 3: Define scraping logic
def extract_data(page, url):
# Human behavior
scraper.human_scroll(page, scrolls=2)
# Extract data (adjust selectors for your site)
title = page.title()
items = page.locator('.item-class').all()
data = []
for item in items:
data.append({
"text": item.text_content(),
"url": url
})
return data
# Step 4: Run scraper
urls = ['https://example.com/page1', 'https://example.com/page2']
results = scraper.scrape_with_rotation(urls, extract_data)
# Step 5: Save results
with open('results.json', 'w') as f:
json.dump(results, f, indent=2)
-
Legal & Ethical Use
- Always respect
robots.txt - Follow website Terms of Service
- Use appropriate rate limiting
- Don't overload servers
- Always respect
-
Proxy Setup
- Replace example proxies with real ones
- Use residential proxies for better success rates
- Rotate proxies to avoid IP bans
-
Tor Browser
- Ensure Tor is running before enabling
use_tor=True - Default SOCKS5 port: 9050
- Install: Tor Project
- Ensure Tor is running before enabling
-
Testing
- Start with
headless=Falseto see what's happening - Test on 1-2 URLs first
- Adjust selectors for your target sites
- Start with
-
Performance
- Human behavior adds delays (realistic but slower)
- Disable in
ScraperConfigif speed is critical - Balance between stealth and performance
Issue: Proxy not working
ProxyServer("http://ip:port") # HTTP ProxyServer("https://ip:port") # HTTPS ProxyServer("socks5://ip:port") # SOCKS5
ProxyServer("http://ip:port", "username", "password")
Issue: Elements not found
page.locator('.your-selector').all()
page.wait_for_selector('.your-selector', timeout=5000)
Issue: Tor not connecting
tor
- Use Multiple Proxies - Rotate through 5-10+ proxies
- Adjust Delays - Match target site's expected behavior
- Monitor Rate Limits - Stay under detection thresholds
- Update Selectors - Websites change, keep selectors current
- Test Thoroughly - Validate on small samples first
This is a template for your projects. Customize and extend as needed!
Use responsibly and ethically. For educational purposes.
⭐ Star this project if it helps you!