Skip to content
WojtekK1902 edited this page Nov 1, 2014 · 7 revisions

FCS platform may be still further developed. Below there are some possible future works presented:

Cloud-dedicated communication

All components of FCS system were prepared and tested on single host and on Vagrant virtual machines environment. Due to some limitations, system was not tested on - still unknown - deployment cloud platform. During installation some aspects of communication between Autoscaling module, Task Server and Crawler should be tested and, if necessary, updated. In future, the autoscaling features and management of Crawling Units are intended to be applied on top cloud providers.

Task Server and Crawler spawning scripts

When concrete cloud environment for deployment will be chosen, new versions of Task Server and Crawler spawning scripts must be provided. Generally, they must be able to run python code from 'web_interface.py' in modules 'fcs.server' and 'fcs.crawler'.

URLs inherited priorities

Every new link which is added into links’ priority queue has default priority. It can be only modified later (directly or indirectly) with feedback. Maybe links extracted from url with higher priority should have higher priority?

Periodic tasks

Once processed page should be crawled once again later. In other words, URL should be added into links' queue once again after some time or amount of links. Feature could be implemented with Huey and Redis tasks’ server (do not confuse with Task Server).

Processing URLs containing a tilde (~)

Python standard library urlparse.urljoin() method cannot handle properly URLs like http://host/~user_name/res (join of http://host/~user_name and /res results with http://host/res). New URL joining method must be implemented.

Auto-adapting statistics

Implemented efficiency estimations use 'url_per_min' parameter, kept in Quota object. It is constant during crawling process, despite even technical limitations connected with database, network and hardware. In the future, 'url_per_min' should be calculated dynamically, taking into consideration recent performance of Crawling Units.

Link database performance optimization

In previous versions of FCS (without feedback feature), Berkeley DB key/value database was used. implementation of user feedback support required new approach to extracted URLs' storage, so graph database Neo4j was applied. Unfortunately, new solution reduced system efficiency. Link database is Task Server's bottleneck and should be optimised if possible.

Robots.txt application

Crawling Unit should be able to work in two modes: respecting robots.txt files and ignoring them. Robots.txt files describes which resources can be accessed by robot and which should not. At this moment The Robots Exclusion Protocol is not supported.

Improved Crawlers' and Task Servers' address management

Actually each new Crawler and Task Server gets address with port one higher that Unit before him. In the future, address should be taken from reusable pool of free ports.

Auto-adjusting Server/Crawler spawn timeout

If during spawning new Server or Crawler waiting for response time is exceeded, new Server/Crawler should be spawned with higher timeout.