Skip to content

Support of multiple docking programs within a single database and implementation of consensus ranking #57

Description

@DrrDom

It seems logical if the same set of ligands is docked by different programs to save docking outputs in the same database. The current implementation requires to make a clean copy of a database, where there are only prepared ligands. This may be somewhat inconvenient in cases if one needs to make consensus ranking. If all outputs will be in the same database, it will be very easy to implement consensus ranking.

This may be a major update which may change the whole structure of a database and the application.

Activity

  1. Feriolet commented on Oct 10, 2025

    @Feriolet
    Contributor

    If we want to do this, then does ALTER TABLE mols ADD COLUMN [insert docking related column] [type] suffice in the respective docking filename? If we want to differentiate different docking tools, we need to have descriptive table column name like vina_docking_score or gnina_mol_block instead of docking_score or mol_block. Of course, this will break backward compatibility for other tools like crem dock and script like the get_sdf_from_dock_db

  2. DrrDom commented on Oct 10, 2025

    @DrrDom
    ContributorAuthor

    This is a possible solution. I do not see the changes in other easydock scripts. like get_sdf_from_dock_db, as a problem. I would like to keep API or at least change it in a compatible way, e.g. add new arguments with appropriate defaults, to not alter integrations. However, this will be hard. I checked and there were about ten easydock functions imported in cremdock, which may a little bit tricky to update.

    Since changes in API is almost unavoidable we may think about deeper refactoring. We can move ligands to a separate table. Docking scores can be stored in another table. A table for docking setup can be created. The issue here is how to keep nice features like automatic docking continuation (e.g. set some variable which docking setup was used the latest and invoke it if it does not differ from input command line arguments or no arguments at all were passed to the script).

    For inspiration I asked chatgpt and claude about their suggestions :)
    https://claude.ai/share/8e2cca64-7870-41f6-803a-adaff2b7875a
    https://chatgpt.com/s/t_68e8c9d0633081919391fa4909be84bb
    Both suggested some superfluous tables, but in general I more like the claude version.

    A minimal set of possible tables:
    ligands - input, processed, stereoisomeric and protonated structures
    docking_programs - this may be a constant table which will list all integrated docking tools irrespective whether a user will use them in a current project
    docking_runs - setup for each docking run and its status
    docking_results - scores, top poses, etc (e.g. source pdb/sdf/mol2 block from which we will retrieve all other poses on-the-fly if requested)
    rescoring_results - the same as docking results + reference to the rescoring program used
    rescoring_runs - the as docking runs but for rescoring programs

    Incompatibilities with external tools we may solve by fixing the version of easydock in other projects and later, if needed, we will refactor them accordingly.

    This are just preliminary thoughts. Maybe some other table can be added to simplify storage and management. This may feel like writing a new program, however, this is mainly a programmer task, scientific tasks were previously solved to some extent.

  3. Feriolet commented on Oct 10, 2025

    @Feriolet
    Contributor

    I see, this solution will definitely rewrite a lot, if not most, of the architecture we currently have. However, this will also allow for highly flexible configurations if we want to pursue this solution (e.g., aside from redocking and rescoring, user will now be able to dock using different protein targets as well and perhaps parallelizing multiple docking software through 1 db) with a cost of significantly more argument to be added.

    For docking_results, should we add a table for each setup? I am concerned that post analysis will be significantly longer if we dump everything into this one table. Or maybe this is too early of an issue to think about it now?

  4. DrrDom commented on Oct 10, 2025

    @DrrDom
    ContributorAuthor

    Good point. With this architecture we will be able to store and use different proteins as well. However, we have to identify somehow identical proteins submitted for different programs which may be used for consensus ranking later. If we will use the same approach with config.yml file I see the only solution to ask a user to provide a name for the protein otherwise it will be extracted from protein file name.

    I would like to keep command line interface as simple as possible, but of course we can add more arguments if needed.

    For docking_results, should we add a table for each setup? I am concerned that post analysis will be significantly longer if we dump everything into this one table. Or maybe this is too early of an issue to think about it now?

    If we refactor the whole architecture we have to take performance into account from early steps to avoid later bottlenecks which could be easy avoided. It will be convenient to use a single table where different runs will have own run_id. This should be also fast to read if this column will be indexed.

    Calculation of consensus ranking most probably will be implemented outside of a database and we will avoid complex operations under DB within a DB engine. If so, we will read relevant rows and columns and all further transformations will be in python code.

    In theory we may think about support of other DB engines using the same interface. This may be a too hard task, however, we may create capacity to implement this later. Claude suggests to create an abstract class DBAdapter and inherit from it.

  5. DrrDom commented on Oct 10, 2025

    @DrrDom
    ContributorAuthor

    The more I think, the more this looks like a completely new program :)

  6. Feriolet commented on Oct 10, 2025

    @Feriolet
    Contributor

    Yeah, I was thinking that we can use the protein filename as a reference to check if the protein used was the same. Though I think it is preferred that the user themselves might give a name instead for better clarity (akin to naming job title so they can manipulate every setup parameter), it may be more convenient for the user to not overwhelm them with so many parameter.

    Regarding run_id, I was also similarly thinking to assign each docking/redocking to this number, so we can save the latest run_id somewhere to preserve the automatic continuation docking feature. Though I do not know if it is better to have the run_id autogenerated through hash, as Claude suggested, or have it user defined.

    Since we want to dump all runs in one table, I agree it is better to have the consensus ranking outside the database.

    Indeed, this feels like an easydock2, since we are reorganising the architecture differently while maintaining its original purpose. We have to map out its full architecture if we decided to pursue this decision.

  7. DrrDom commented on Oct 10, 2025

    @DrrDom
    ContributorAuthor

    Yeah, I was thinking that we can use the protein filename as a reference to check if the protein used was the same. Though I think it is preferred that the user themselves might give a name instead for better clarity (akin to naming job title so they can manipulate every setup parameter), it may be more convenient for the user to not overwhelm them with so many parameter.

    Making an option to set a name of a protein is necessary also because different programs will require different inputs and they may have own preparation pipelines with other naming requirements. However, this will be optional and it will may be easy to change the protein name (alias) in DB later but direct editing.

    Regarding run_id, I was also similarly thinking to assign each docking/redocking to this number, so we can save the latest run_id somewhere to preserve the automatic continuation docking feature. Though I do not know if it is better to have the run_id autogenerated through hash, as Claude suggested, or have it user defined.

    Definitely a user should not set run_id manually. This is an internal identifier. The idea with a hash looks interesting. If we will use config, protein, girdbox and other files as input, we can store them inside DB and hash all files together with input arguments (including default values) to get a single unique identifier of run settings. However, I suspect there may be some difficulties, because one byte change will change the hash.

    Since we want to dump all runs in one table, I agree it is better to have the consensus ranking outside the database.

    Some consensus scores can be computed by DB, but some may be more tricky or slow. Therefore, in general this can be done outside of DB.

  8. Feriolet commented on Oct 11, 2025

    @Feriolet
    Contributor

    Making an option to set a name of a protein is necessary also because different programs will require different inputs and they may have own preparation pipelines with other naming requirements. However, this will be optional and it will may be easy to change the protein name (alias) in DB later but direct editing.

    I see, that should be fine then.

    Definitely a user should not set run_id manually. This is an internal identifier. The idea with a hash looks interesting. If we will use config, protein, girdbox and other files as input, we can store them inside DB and hash all files together with input arguments (including default values) to get a single unique identifier of run settings. However, I suspect there may be some difficulties, because one byte change will change the hash.

    Yeah, this one might be tricky because each user may have different preference that makes consistent hash difficult. If we hash based on config content (i.e., hashing [protein_name, protein_fname, grid, etc], user can use the same protein with different directories which alters hash value. If we hash based on file content (i.e., hashing [protein_name, PDB content, grid content, etc], accidentally changing 1 byte of protein will also create a new hash value.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions