haddock.libs.libligand module

Ligand topology utilities.

haddock.libs.libligand.extract_ligand(pdb_file: str | Path, resnames: list[str], dest: str | Path) → Path[source]

Write a PDB with a single copy of each of the given residue names.

prodrg can only handle a single small molecule. When the ligand is part of a larger system (e.g. a protein/ligand complex), the surrounding macromolecule must be stripped out before prodrg is run, otherwise it fails to generate topology and parameter files.

When a ligand is present in several copies (same residue name, different residue numbers/chains), all copies share the same topology, so only the first copy of each residue name is kept to avoid feeding prodrg redundant (and potentially too many) atoms.

Atoms that CNS mislabelled as a two-letter metal element (e.g. a phosphate PB written with element Pb) are corrected back to their real single-letter element so prodrg can parametrise them.

Parameters:
  • pdb_file (FilePath) – Path to the input PDB file (may contain a full system).

  • resnames (list[str]) – Residue names to keep (typically the unknown ligand residues).

  • dest (FilePath) – Path to the PDB file to write with only the selected residues.

Returns:

Path – Path to the written ligand-only PDB file (dest).

haddock.libs.libligand.identify_unknown_hetatms(pdb_file: str | Path) → list[str][source]

Return residue names in a PDB that are not in the supported residues.

Parameters:

pdb_file (FilePath) – Path to the PDB file to inspect.

Returns:

list[str] – Unique residue names not found in the supported residues set, in order of first appearance.

haddock.libs.libligand.run_prodrg(pdb_file: str | Path, output_dir: str | Path, ligand_resnames: list[str] | None = None) → tuple[Path, Path][source]

Run prodrg on a ligand PDB and write CNS topology and parameter files.

The resulting files are written to output_dir named after the input PDB stem.

prodrg can only process a single small molecule at a time. When ligand_resnames lists several distinct ligands, prodrg is run once per ligand (on a single extracted copy of each) and the resulting topology and parameter files are concatenated into one .top and one .param. Each ligand is given a unique CNSSEP character so that prodrg generates non-overlapping atom type names across the concatenated files; the characters used are also chosen to avoid the separators already present in the built-in cofactors.top topology, so a single auto-generated ligand likewise cannot clash with the cofactor atom types.

Parameters:
  • pdb_file (FilePath) – Path to the ligand PDB file. It may also be a full system (e.g. a protein/ligand complex), in which case ligand_resnames must be provided so that only the ligand residues are passed to prodrg.

  • output_dir (FilePath) – Directory where the named .top and .param files are written.

  • ligand_resnames (Optional[list[str]]) – Residue names of the ligand(s) to extract before running prodrg. When provided, each distinct ligand is extracted (a single copy) and passed to prodrg separately, stripping out the rest of the system (e.g. the protein). When None, the whole pdb_file is passed to prodrg unchanged.

Returns:

tuple[Path, Path] – Paths to the written <stem>_prodrg.top and <stem>_prodrg.param files.

Raises:

RuntimeError – If prodrg exits with a non-zero return code or the expected output files are not created.