Isomer counting¶
Single-group isomer counting (Burnside's lemma) and enumeration helpers used by enumerate_isomers, and the legacy FG-FG statistics: cage_isomer_builder.utils.util.
count_unique_isomers_k ¶
count_unique_isomers_k(transformation_library, n_linkers, n_fg_slots, fg_per_linker=4, active_per_linker=1)
Number of symmetry-unique isomers with exactly active_per_linker active
slots on every linker, from Burnside's lemma (count_unique_isomers is the
case active_per_linker=1).
Burnside's lemma: the number of orbits of a group G acting on a set is
n_unique = (1/|G|) * sum over g in G of |Fix(g)|
The set here is every choice of k active slots per linker, so the identity
fixes C(fg_per_linker, k) ** n_linkers descriptors; every other |Fix(g)|
comes from g's cycles (see _fixed_isomer_count_k). The average is a whole
number for a group action, and the division is checked.
Checks for a result: k = 0 and k = fg_per_linker give 1; k and fg_per_linker - k give the same count; k = 1 equals count_unique_isomers.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
transformation_library
|
list of list of int
|
Symmetry operations as slot permutations (operation-major), as for count_unique_isomers. |
required |
n_linkers
|
int
|
Number of linkers. |
required |
n_fg_slots
|
int
|
Total number of slots (n_linkers * fg_per_linker). |
required |
fg_per_linker
|
int
|
Slots per linker. |
4
|
active_per_linker
|
int
|
Active slots on every linker, between 0 and fg_per_linker. |
1
|
Returns:
| Type | Description |
|---|---|
int
|
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If active_per_linker is out of range, or the Burnside average is not a whole number (the operations are not a group on these slots). |
Examples:
>>> from cage_isomer_builder.utils.util import count_unique_isomers_k
>>> swap = [[2, 3, 0, 1]] # two 2-slot linkers swapped by symmetry
>>> count_unique_isomers_k(swap, n_linkers=2, n_fg_slots=4, fg_per_linker=2, active_per_linker=1)
3
Source code in cage_isomer_builder/utils/util.py
149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 | |
count_unique_isomers ¶
Number of symmetry-unique isomers with one active slot per linker, from Burnside's lemma:
n_unique = (1/|G|) * sum over g in G of |Fix(g)|
where |Fix(g)| is the number of descriptors left unchanged by operation g (computed exactly from g's cycles, see _fixed_isomer_count) and |G| counts the identity.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
transformation_library
|
list of list of int
|
Symmetry operations as slot permutations: |
required |
n_linkers
|
int
|
Number of linkers; there are |
required |
n_fg_slots
|
int
|
Total number of slots (n_linkers * fg_per_linker). |
required |
fg_per_linker
|
int
|
Slots per linker (4 for a phenylene, 8 for a biphenylene, ...). |
4
|
Returns:
| Type | Description |
|---|---|
int
|
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If n_linkers is above 24 (the table becomes impractical; use
|
Examples:
>>> from cage_isomer_builder.utils.util import count_unique_isomers
>>> count_unique_isomers([[2, 3, 0, 1]], n_linkers=2, n_fg_slots=4, fg_per_linker=2)
3
Source code in cage_isomer_builder/utils/util.py
fg2fg_distance_count ¶
Legacy FG-FG pair counts, kept for compatibility. CageBuilder.get_statistics
does not use it: its endo/exo rule (the slots of each linker closest to the
origin are endo) assumes the structure is centred at the origin and depends
on sub-Angstrom differences when a ring lies face-on to the pore, and its
distances are not minimum-image for periodic cells. Use
CageBuilder.get_statistics or
:func:cage_isomer_builder.utils.distributions.pair_descriptor instead.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
ase_atom
|
Atoms
|
The site atoms, in slot order. |
required |
fg_anchor_indices
|
list of int
|
Slot indices (one entry per site). |
required |
fg_per_linker
|
int
|
|
4
|
Returns:
| Name | Type | Description |
|---|---|---|
distance_keys |
list of float
|
Sorted distinct distances between slots on different linkers. |
total_count |
list of int
|
Number of slot pairs at each distance. |
inner_inner_count, inner_outer_count, outer_outer_count : list of int
|
The same, split by the legacy endo/exo rule. |
Examples:
>>> from ase import Atoms
>>> from cage_isomer_builder.utils.util import fg2fg_distance_count
>>> slots = Atoms("X4", positions=[[0, 0, 0], [1, 0, 0], [5, 0, 0], [6, 0, 0]])
>>> keys, total, *_ = fg2fg_distance_count(slots, [0, 1, 2, 3], fg_per_linker=2)
>>> [float(k) for k in keys], total # pairs on different linkers only
([4.0, 5.0, 6.0], [1, 2, 1])
Source code in cage_isomer_builder/utils/util.py
305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 | |
kernel_density_estimation ¶
Smoothed version of a distance histogram (Gaussian kernel density estimate).
Each distance is repeated by its count, a Gaussian KDE (bandwidth 0.1 Angstrom) is fitted, and the curve is scaled so that its values add up to the total count, so it can be drawn on the same axis as the histogram.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
distance_values
|
list of float
|
Sorted distinct distances (Angstrom). |
required |
pair_counts
|
list of int
|
Number of pairs at each distance. |
required |
n_bins
|
int
|
Number of evaluation points, from 0 to 1.05 * max(distance_values). |
required |
Returns:
| Name | Type | Description |
|---|---|---|
x_eval |
(ndarray, shape(n_bins, 1))
|
Evaluation grid (Angstrom). |
y_density |
(ndarray, shape(n_bins))
|
Curve values, adding up to the total pair count. |
Examples:
>>> from cage_isomer_builder.utils.util import kernel_density_estimation
>>> x, y = kernel_density_estimation([4.0, 5.0], [1, 2], n_bins=200)
>>> x.shape, round(float(y.sum()), 6) # rescaled to the pair count
((200, 1), 3.0)
Source code in cage_isomer_builder/utils/util.py
convert_base4_to_base10 ¶
Convert a base-4 isomer descriptor into an integer index.
Each element of the isomer descriptor encodes the rotational state (0-3) of a functional group anchor on a given linker. The decimal index is computed as a base-4 number where each digit corresponds to the rotational state of one linker.
Schematically for n_linkers = 4: isomer_descriptor = [slot_0, slot_1, slot_2, slot_3] decimal_index = r0 * 4^3 + r1 * 4^2 + r2 * 4^1 + r3 * 4^0 where r_i = isomer_descriptor[i] % 4 (rotational state of linker i)
This is the inverse of convert_base10_to_base4.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
isomer_descriptor
|
list of int
|
Base-4 isomer descriptor of length n_linkers, where each element is the global FG anchor slot index active on that linker. The rotational state of linker i is recovered as isomer_descriptor[i] % 4, giving a value in range 0-3. e.g. [0, 5, 9, 14] for a cage with 4 linkers. |
required |
n_linkers
|
int
|
Number of linkers in the cage. Determines the number of base-4 digits and the positional weight of each digit. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
decimal_index |
int
|
Index of the descriptor, from 0 to fg_per_linker*n_linkers - 1, e.g. [0, 1, 0, 2] -> 064 + 116 + 04 + 2*1 = 18. |
Examples:
>>> from cage_isomer_builder.utils.util import convert_base4_to_base10
>>> convert_base4_to_base10([0, 5, 8, 14], n_linkers=4) # local slots 0, 1, 0, 2
18
Source code in cage_isomer_builder/utils/util.py
convert_base10_to_base4 ¶
Convert a base-10 (decimal) integer index into a base-4 isomer descriptor.
The decimal index is decomposed into base-4 digits by successive division, where each digit represents the rotational state (0-3) of one linker. The digits are then converted into global FG anchor slot indices by adding the linker offset (linker_index * 4).
Schematically for n_linkers = 4: decimal_index = 18 base-4 digits = [0, 1, 0, 2] (most to least significant) isomer_descriptor = [04+0, 14+1, 24+0, 34+2] = [0, 5, 8, 14]
This is the inverse of convert_base4_to_base10.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
decimal_index
|
int
|
Index of the descriptor, from 0 to fg_per_linker**n_linkers - 1. |
required |
n_linkers
|
int
|
Number of linkers in the cage. Determines the number of base-4 digits and the global slot index offset for each linker. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
isomer_descriptor |
list of int
|
Base-4 isomer descriptor of length n_linkers, where each element is the global FG anchor slot index active on that linker. The rotational state of linker i is recovered as isomer_descriptor[i] % 4, giving a value in range 0-3. e.g. decimal_index=18, n_linkers=4 -> [0, 5, 8, 14] |
Examples:
>>> from cage_isomer_builder.utils.util import convert_base10_to_base4
>>> convert_base10_to_base4(18, n_linkers=4)
[0, 5, 8, 14]
Source code in cage_isomer_builder/utils/util.py
iter_unique_isomers ¶
Yield the symmetry-unique isomers with one active slot per linker as (decimal_index, descriptor) pairs, in ascending index order.
A candidate is kept if it is the smallest index of its orbit (see
_is_canonical_isomer). Each candidate is checked against its own orbit, so
memory use does not grow with the number of linkers; the run time grows with
fg_per_linker ** n_linkers unless limit stops early. Canonical
representatives occur about once per |G| indices, so a limit is reached
quickly even when the full walk is not feasible.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
tl
|
list of tuple
|
Slot-major transformation library: |
required |
n_linkers
|
int
|
|
required |
fg_per_linker
|
int
|
|
4
|
limit
|
int
|
Stop after this many isomers. Default: walk every raw descriptor. |
None
|
Yields:
| Name | Type | Description |
|---|---|---|
i_isomer |
int
|
Decimal index of the isomer (see convert_base10_to_base4). |
descriptor |
list of int
|
The active slot of each linker. |
Examples:
>>> from cage_isomer_builder.utils.util import iter_unique_isomers
>>> library = [(2,), (3,), (0,), (1,)] # slot-major: one swap of two linkers
>>> [d for _, d in iter_unique_isomers(library, n_linkers=2, fg_per_linker=2)]
[[0, 2], [0, 3], [1, 3]]
Source code in cage_isomer_builder/utils/util.py
is_higher_greater ¶
True if the first isomer descriptor is lexicographically greater than the second.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
isomer_descriptor_1
|
list of int
|
First base-4 isomer descriptor of length n_linkers, where each element is the global FG anchor slot index active on that linker. e.g. [0, 5, 9, 14] |
required |
isomer_descriptor_2
|
list of int
|
Second base-4 isomer descriptor of length n_linkers to compare against. Must be the same length as isomer_descriptor_1. e.g. [0, 5, 9, 10] |
required |
Returns:
| Type | Description |
|---|---|
bool
|
True if isomer_descriptor_1 is lexicographically greater than isomer_descriptor_2, False otherwise (including when equal). |
Examples:
>>> from cage_isomer_builder.utils.util import is_higher_greater
>>> is_higher_greater([0, 5, 9, 14], [0, 5, 9, 10]), is_higher_greater([0, 4], [0, 4])
(True, False)