A professional web application for finding business contact information from multiple sources including Meta Business Pages, Google Maps, Google My Business, TikTok, and direct website scraping.
- Multi-Source Search: Search across Meta, Google Maps, Google My Business, and TikTok
- Website Scraper: Extract contact info directly from business websites
- Results Management: View, filter, and manage scraped contacts
- Export Options: Export to CSV and JSON formats
- Compliance Dashboard: Legal notices and best practices
- Node.js Scraper: Puppeteer-based scraper for JavaScript-heavy sites
- Python Scraper: BeautifulSoup + Playwright for dynamic content
- REST API: Full API for programmatic access
- Ethical Scraping: Respects robots.txt, rate limiting, proper user agents
├── src/ # React frontend
│ ├── App.tsx # Main application
│ ├── WebsiteScraperView.tsx # Website scraper UI
│ ├── types.ts # TypeScript types
│ └── mockData.ts # Demo data
├── backend/ # Backend scraping services
│ ├── scraper-service.js # Node.js scraper (Puppeteer)
│ ├── scraper.py # Python scraper (BeautifulSoup)
│ ├── package.json # Node.js dependencies
│ └── requirements.txt # Python dependencies
└── SCRAPING_GUIDE.md # Comprehensive scraping guide
The frontend is already built and ready to use. To run in development mode:
npm install
npm run devThe app will be available at http://localhost:5173
cd backend
npm install
npm startThe scraper service will run on http://localhost:3001
Endpoints:
POST /api/scrape- Scrape single URLPOST /api/scrape/batch- Scrape multiple URLsPOST /api/extract-emails- Extract emails from textGET /api/health- Health check
cd backend
pip install -r requirements.txt
playwright install chromium # For dynamic sites
python scraper.py --serveThe API server will run on http://localhost:8000
Endpoints:
POST /api/scrape?url=<url>&deep=<true|false>- Scrape URLPOST /api/scrape/batch- Scrape multiple URLsPOST /api/extract-emails- Extract emails from textGET /api/health- Health check
- Go to the Search tab
- Enter business name/keyword and location
- Select data sources (Meta, Google Maps, GMB, TikTok)
- Click "Start Search"
- View results in the Results tab
- Export from the Export tab
- Go to the Web Scraper tab
- Enter one or more website URLs (one per line)
- Optionally enable "Deep Scrape" to follow contact links
- Click "Start Scraping"
- View extracted emails, phones, addresses, and social links
Note: Make sure the backend scraper service is running for real scraping. Without it, the frontend will show demo results.
- Go to the Export tab
- Choose format (CSV or JSON)
- Download your contacts
In the Settings tab, you can configure:
- API keys for Meta, Google Maps, and TikTok
- Request delays and rate limiting
- Concurrent requests
- Email validation options
Node.js Scraper (scraper-service.js):
const CONFIG = {
requestDelay: 2000, // Delay between requests (ms)
timeout: 30000, // Page load timeout (ms)
maxRetries: 3, // Max retry attempts
userAgent: 'ContactHarvest/1.0',
maxConcurrent: 3, // Max concurrent scrapes
};Python Scraper (scraper.py):
@dataclass
class ScraperConfig:
request_delay: tuple = (2, 5) # Random delay range (seconds)
timeout: int = 30
max_retries: int = 3
max_pages: int = 5
respect_robots: bool = True
validate_emails: bool = True-
Meta Graph API - Facebook & Instagram business data
-
Google Places API - Business location data
-
Google Business Profile API - GMB data
-
TikTok for Developers - TikTok business data
For websites without APIs, use the built-in scraper:
- Extracts emails, phones, addresses
- Follows contact/about pages (deep scrape)
- Respects robots.txt
- Rate limited to avoid overwhelming servers
This tool is for educational and demonstration purposes. Before using any web scraping tool:
- Check Platform ToS: Each platform (Meta, Google, TikTok) has specific terms
- Use Official APIs: Always prefer official APIs over scraping
- Respect robots.txt: Check before scraping any website
- Rate Limiting: Don't overwhelm servers with requests
- Privacy Laws: Comply with GDPR, CCPA, and other regulations
- Opt-Out: Provide mechanisms for businesses to opt out
✅ DO:
- Use official APIs when available
- Check robots.txt before scraping
- Respect rate limits (2-5 second delays)
- Only scrape publicly available business information
- Provide opt-out mechanisms
- Comply with privacy laws
❌ DON'T:
- Scrape behind login walls without permission
- Ignore robots.txt restrictions
- Overwhelm servers with requests
- Scrape personal (non-business) information
- Violate platform Terms of Service
- Use data for spam or harassment
- Fast and lightweight
- Works for simple HTML pages
- Tools:
requests+BeautifulSoup(Python),axios+cheerio(Node.js)
- Required for JavaScript-heavy sites
- Slower but more powerful
- Tools:
Puppeteer(Node.js),Playwright(Python/Node.js)
- Regex pattern matching
- Mailto link detection
- Schema.org structured data
- Contact page crawling
See SCRAPING_GUIDE.md for detailed technical documentation.
If the frontend shows "Backend service not available":
-
Make sure the backend is running:
cd backend npm start # or python scraper.py --serve
-
Check the backend URL in the Web Scraper tab matches your backend port
-
Verify CORS is enabled (should be by default)
Common issues:
- Blocked by robots.txt: The site doesn't allow scraping
- Timeout: Site is slow or blocking automated access
- Invalid URL: Check URL format (include https://)
- Rate limited: Too many requests, increase delay
If the backend is not running, the frontend will show demo results. This is normal for testing the UI.
curl -X POST http://localhost:3001/api/scrape \
-H "Content-Type: application/json" \
-d '{"url": "https://example-business.com", "deep": true}'curl -X POST http://localhost:3001/api/scrape/batch \
-H "Content-Type: application/json" \
-d '{"urls": ["https://business1.com", "https://business2.com"]}'# Scrape single URL
python scraper.py --url https://example-business.com
# Deep scrape
python scraper.py --url https://example-business.com --deep
# Batch scrape from file
python scraper.py --batch urls.txt --output results.jsonFor production use:
- Set up proper API keys for Meta, Google, TikTok
- Use a queue system (Redis + Bull) for managing scraping jobs
- Add authentication to your backend API
- Implement database storage (PostgreSQL recommended)
- Add monitoring and logging
- Use rotating proxies to avoid IP bans
- Implement proper error handling
- Add email validation before storing contacts
- SCRAPING_GUIDE.md - Comprehensive scraping guide
- Puppeteer Documentation
- Playwright Documentation
- BeautifulSoup Documentation
- Meta Graph API
- Google Places API
This is a demonstration project. For production use, ensure compliance with all applicable laws and platform terms of service.
This project is for educational purposes. Always comply with platform terms of service and applicable laws when scraping websites.