Guest matching¶
Guest conformer descriptors and host-guest matching: cage_isomer_builder.utils.matching.
Describe a guest the same way as the host and match the two (Objective 3).
Guest descriptor
A flexible guest is a set of conformers from a GA/PES scan, MD snapshots or
any conformer generator; this module does not generate them. Each
conformer's feature sites (:func:cage_isomer_builder.utils.features.detect_features)
give a pair descriptor; the guest descriptor is their weighted average,
with Boltzmann weights from conformer energies or uniform weights (e.g. for
MD snapshots, already Boltzmann-distributed by the simulation).
Matching
A guest binds where the host offers complementary features at the same separations as the guest's own: a guest donor-acceptor pair at distance d fits a host acceptor-donor pair at about d. For every guest pair type (a, b) the host pair types (a', b') with a' complementary to a and b' to b are pooled, and the two distance distributions are compared after normalising each to unit area (overlap = integral of min(p, q), 0 to 1). The total score is the average over guest pair types, weighted by the guest's pair weights, so it is also between 0 and 1.
Default complementarity (:data:COMPLEMENTARY):
- guest donor -> host acceptor
- guest acceptor -> host donor, host open metal site
- guest pi -> host pi
- guest fg -> host fg (same label), for functional-group fingerprints
This is a screening score, a fast way to rank pores, isomers and guests before computing host-guest energies; it is not an energy. The distance compared is between features of the same molecule, so contact offsets (e.g. the H...O distance) cancel to first order but are not modelled.
MatchResult
dataclass
¶
Attributes:
| Name | Type | Description |
|---|---|---|
score |
float
|
Weighted mean overlap over guest pair types, 0 to 1. |
per_pair |
dict
|
|
unmatched |
list
|
Guest pair types for which the host has no complementary pairs (they score 0). |
defined |
bool
|
False when the guest has no feature pairs at all (e.g. a single
aromatic ring and nothing else): pair distributions then say nothing
and |
Examples:
>>> from ase.build import molecule
>>> from cage_isomer_builder.utils.matching import conformer_descriptor, match_descriptors
>>> water = conformer_descriptor([molecule("H2O")])
>>> result = match_descriptors(water, water)
>>> result.defined, sorted(result.per_pair)
(True, [('acceptor', 'donor'), ('donor', 'donor')])
conformer_descriptor ¶
conformer_descriptor(conformers, energies=None, temperature=298.15, unit='eV', weights=None, kinds=GUEST_KINDS, by='kind')
Pair descriptor of a guest averaged over its conformers.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
conformers
|
sequence of ase.Atoms
|
All the same molecule (any orientation). |
required |
energies
|
sequence of float
|
Conformer energies, for Boltzmann weights at |
None
|
weights
|
sequence of float
|
Explicit weights instead (normalised). Default: uniform. |
None
|
kinds
|
sequence of str
|
Feature kinds to use. |
GUEST_KINDS
|
by
|
str
|
Pair typing (see :func: |
'kind'
|
Returns:
| Type | Description |
|---|---|
Descriptor
|
Expected pairs per molecule. |
Examples:
>>> from ase.build import molecule
>>> from cage_isomer_builder.utils.matching import conformer_descriptor
>>> water = molecule("H2O")
>>> bent = water.copy()
>>> bent.positions[1] *= 1.05 # a second, slightly different conformer
>>> desc = conformer_descriptor([water, bent], energies=[0.0, 0.02])
>>> [round(w, 3) for w in desc.info["weights"]] # Boltzmann weights at 298 K
[0.685, 0.315]
>>> desc.total() # 2 acceptor-donor + 1 donor-donor per molecule
3.0
Source code in cage_isomer_builder/utils/matching.py
match_descriptors ¶
Score how well a guest's feature-pair distances fit the host's complementary feature-pair distances (see the module docstring).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
host
|
Descriptor
|
Typed with the same |
required |
guest
|
Descriptor
|
Typed with the same |
required |
grid
|
array - like
|
Distance grid (Angstrom). Default 0 to 30 in 0.02 steps. |
None
|
bandwidth
|
float
|
KDE width (Angstrom), wider than for host-only statistics because a binding pocket tolerates some mismatch. |
0.25
|
metric
|
(overlap, bhattacharyya)
|
|
"overlap"
|
complementarity
|
dict
|
Overrides :data: |
None
|
Returns:
| Type | Description |
|---|---|
MatchResult
|
|
Examples:
>>> import numpy as np
>>> from cage_isomer_builder.utils.distributions import Descriptor, PairSet
>>> from cage_isomer_builder.utils.matching import match_descriptors
>>> def pairs(key, d):
... out = Descriptor()
... out.add(key, PairSet(np.array([d]), np.array([np.nan]), np.array([1.0])))
... return out
>>> guest = pairs(("acceptor", "donor"), 5.0) # guest donor and acceptor 5 A apart
>>> round(match_descriptors(pairs(("acceptor", "donor"), 5.0), guest).score, 3)
1.0
>>> round(match_descriptors(pairs(("acceptor", "donor"), 7.0), guest).score, 3)
0.0
>>> match_descriptors(pairs(("pi", "pi"), 5.0), guest).unmatched # nothing complementary
[('acceptor', 'donor')]
Source code in cage_isomer_builder/utils/matching.py
rank_hosts ¶
[(name, MatchResult), ...] best first, for {name: Descriptor}
hosts (e.g. pores, isomers or defect variants).
Examples:
>>> import numpy as np
>>> from cage_isomer_builder.utils.distributions import Descriptor, PairSet
>>> from cage_isomer_builder.utils.matching import rank_hosts
>>> def pairs(d):
... out = Descriptor()
... out.add(("acceptor", "donor"), PairSet(np.array([d]), np.array([np.nan]), np.array([1.0])))
... return out
>>> ranked = rank_hosts({"small pore": pairs(4.0), "matching pore": pairs(5.0)}, pairs(5.0))
>>> [name for name, _ in ranked]
['matching pore', 'small pore']