> For the complete documentation index, see [llms.txt](https://ersilia.gitbook.io/ersilia-workshops/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://ersilia.gitbook.io/ersilia-workshops/ai4dd-latam/sessions/monday-afternoon-breakout.md).

# Monday afternoon (breakout)

Hands-on exercise using AI/ML models from the Ersilia Model Hub

In this session, we will do a gropu exercise using AI/ML models from the Ersilia Model Hub to predict compound properties and their potential antimicrobial activity.

We will split into groups. Each group will be assigned a pathogen (*Acinetobacter baumannii* or *Staphylococcus aureus*) and a screening compound library, including \~1000 small molecules from the [ChemDiv Anti-Infective Library](https://www.chemdiv.com/catalog/focused-and-targeted-libraries/anti-infective-library/). At the end of the activity, each group will be asked to present their compounds of choice.

## Groups

Below is the pathogen of interest assigned to each group.

#### *Acinetobacter baumannii*

* <mark style="background-color:green;">Green</mark>
* <mark style="background-color:yellow;">Yellow</mark>
* <mark style="background-color:blue;">Blue</mark>

#### *Staphylococcus aureus*

* <mark style="background-color:orange;">Orange</mark>
* <mark style="background-color:$danger;">Pink</mark>

## Steps

1. Group creation (nominate a scribe)
2. Download the compound library corresponding to your groups from [here](https://drive.google.com/drive/folders/1BAuc0doyocoPaL-wtSAdB-Sr1q73uDdM?usp=drive_link).
3. Look at the table [below](#relevant-models-from-the-ersilia-model-hub) and select up to 4 models relevant for your task.
4. Discuss the publications related to each model and note down what kind of model it is (classifier, regressor), which output will it give and how to interpret it.
5. Go to the [Ersilia GUI](https://hub.ersilia.io/) and run evaluations for the selected models.
6. Download all CSV files into a working directory. In case a model evaluation fails, feel free to download precalculations from [here](https://drive.google.com/drive/folders/1x5kqLuPXm_GZZnOAOTPs9R8lVKr5j9GS?usp=drive_link).
7. Use Excel to concatenate the CSV files and [this app](https://pubchem.ncbi.nlm.nih.gov//edit3/index.html) to visualize compounds.
8. Select up to 5 compounds.
9. Make a **short** slide deck. We suggest 3 slides:
   * Selected models and rationale. Why is your pathogen relevant? Which models and columns did you choose, and why?
   * Overview of the screening results. Were the results as expected?
   * Selected compounds and rationale. Which compounds did you choose, and why? Did you identify an optimal compound? What would you do next?
10. Present! Send your slides to <miquel@ersilia.io>. Name your file: ai\_workshop\_green.pptx, ai\_workshop\_yellow\.pptx, etc.

## Materials

* [Slide deck](https://drive.google.com/file/d/1CbvhGbcWVnfmXorg8IT-uqZfInP9ovTh/view?usp=drive_link)
* [Compound libraries](https://drive.google.com/drive/folders/1HUotCn4NQCW8qtftsmbeP3RHoVL-SNKL?usp=drive_link)
* [Precalculated results](https://drive.google.com/drive/folders/1x5kqLuPXm_GZZnOAOTPs9R8lVKr5j9GS?usp=drive_link)
* [Ersilia GUI](https://hub.ersilia.io)
* [Draw your molecules online](https://pubchem.ncbi.nlm.nih.gov//edit3/index.html)
* [Publicaciones](https://drive.google.com/drive/folders/1JtfHENYN3S_iCvohNKoH2KBTOkih70yG?usp=drive_link)
* [Participant presentations](https://drive.google.com/drive/folders/1v_V5M4JLigpVKmqxSrmMXntbukfsOqN2?usp=drive_link)

## Relevant models from the Ersilia Model Hub

Below is a list of relevant models from the Ersilia Model Hub. Click the model identifier to access the model repository, where you can find more information about each of the models, including a description of the columns.

{% hint style="info" %}
To view information for each of the columns of a given molecule, you have to (a) access the model repository using the link in the table below, and (b) visit the `/model/framework/columns/run_columns.csv` .
{% endhint %}

<table><thead><tr><th width="97.46875">Model ID</th><th width="140.46875">Slug</th><th width="259.05078125">Title</th><th>Columns</th><th width="658.59765625">Description</th><th data-hidden>Link</th></tr></thead><tbody><tr><td><a href="https://github.com/ersilia-os/eos3804"><strong>eos3804</strong></a></td><td>chemprop-abaumannii</td><td>Inhibition of Acinetobacter baumannii growth</td><td><a href="https://github.com/ersilia-os/eos3804/blob/main/model/framework/columns/run_columns.csv">1 column</a></td><td>This model is a Chemprop neural network trained with a growth inhibition dataset. Authors screened ~7,500 molecules for those that inhibited the growth of A. baumannii in vitro. They discovered abaucin, an antibacterial compound with narrow-spectrum activity against A. baumannii.</td><td><a href="https://github.com/ersilia-os/eos3804">https://github.com/ersilia-os/eos3804</a></td></tr><tr><td><a href="https://github.com/ersilia-os/eos42ez"><strong>eos42ez</strong></a></td><td>antibiotics-ai-cytotox</td><td>Human cytotoxicity endpoints</td><td><a href="https://github.com/ersilia-os/eos42ez/blob/main/model/framework/columns/run_columns.csv">3 columns</a></td><td>The authors tested the dataset of 39312 compounds used to train the antibiotics-ai model (eos18ie) against several cytotoxicity endpoints; human liver carcinoma cells (HepG2), human primary skeletal muscle cells (HSkMCs) and human lung fibroblast cells (IMR-90). Cellular viability was measured after 20133 days of treatment with each compound at 10 μM and activities were binarized using a 90% cell viability cut-off. 341 (8.5%), 490 (3.8%) and 447 (8.8%) compounds classified as cytotoxic for HepG2 cells, HSk-MCs and IMR-90 cells</td><td><a href="https://github.com/ersilia-os/eos42ez">https://github.com/ersilia-os/eos42ez</a></td></tr><tr><td><a href="https://github.com/ersilia-os/eos2db3"><strong>eos2db3</strong></a></td><td>chemical-space-projections-chemdiv</td><td>Chemical space 2D projections against ChemDiv</td><td><a href="https://github.com/ersilia-os/eos2db3/blob/main/model/framework/columns/run_columns.csv">8 columns</a></td><td>This tool performs PCA, UMAP and tSNE projections taking a 100k ChemDiv diversity set as a chemical space of reference. The Ersilia Compound Embeddings are used as descriptors. Four PCA components and two UMAP and tSNE components are returned.</td><td><a href="https://github.com/ersilia-os/eos2db3">https://github.com/ersilia-os/eos2db3</a></td></tr><tr><td><a href="https://github.com/ersilia-os/eos18ie"><strong>eos18ie</strong></a></td><td>antibiotics-ai-saureus</td><td>Antibiotic activity prediction against Staphylococcus aureus</td><td><a href="https://github.com/ersilia-os/eos18ie/blob/main/model/framework/columns/run_columns.csv">1 column</a></td><td>The authors use a mid-size dataset (more than 30k compounds) to train an explainable graph-based model to identify potential antibiotics with low cytotoxicity. The model uses a substructure-based approach to explore the chemical space. Using this method, they were able to screen 283 compounds and identify a candidate active against methicillin-resistant S. aureus (MRSA) and vancomycin-resistant enterococci.</td><td><a href="https://github.com/ersilia-os/eos18ie">https://github.com/ersilia-os/eos18ie</a></td></tr><tr><td><a href="https://github.com/ersilia-os/eos37l0"><strong>eos37l0</strong></a></td><td>chembl-kpneumoniae</td><td>Klebsiella pneumoniae activity prediction</td><td><a href="https://github.com/ersilia-os/eos37l0/blob/main/model/framework/columns/run_columns.csv">22 columns</a></td><td>Klebsiella pneumoniae activity prediction based on phenotypic ChEMBL data. Each column corresponds to a specific bioactivity dataset derived from ChEMBL, encompassing multiple assays and binarization cut-offs. The global consensus score summarizes the probability of being active. Model developed by Ersilia.</td><td><a href="https://github.com/ersilia-os/eos37l0">https://github.com/ersilia-os/eos37l0</a></td></tr><tr><td><a href="https://github.com/ersilia-os/eos5dti"><strong>eos5dti</strong></a></td><td>chembl-abaumannii</td><td>Acinetobacter baumannii activity prediction</td><td><a href="https://github.com/ersilia-os/eos5dti/blob/main/model/framework/columns/run_columns.csv">26 columns</a></td><td>Acinetobacter baumannii activity prediction based on phenotypic ChEMBL data. Each column corresponds to a specific bioactivity dataset derived from ChEMBL, encompassing multiple assays and binarization cut-offs. The global consensus score summarizes the probability of being active. Model developed by Ersilia.</td><td><a href="https://github.com/ersilia-os/eos5dti">https://github.com/ersilia-os/eos5dti</a></td></tr><tr><td><a href="https://github.com/ersilia-os/eos2m0f"><strong>eos2m0f</strong></a></td><td>chembl-saureus</td><td>Staphylococcus aureus activity prediction</td><td><a href="https://github.com/ersilia-os/eos2m0f/blob/main/model/framework/columns/run_columns.csv">51 columns</a></td><td>Staphylococcus aureus activity prediction based on phenotypic ChEMBL data. Each column corresponds to a specific bioactivity dataset derived from ChEMBL, encompassing multiple assays and binarization cut-offs. The global consensus score summarizes the probability of being active. Model developed by Ersilia.</td><td><a href="https://github.com/ersilia-os/eos2m0f">https://github.com/ersilia-os/eos2m0f</a></td></tr><tr><td><a href="https://github.com/ersilia-os/eos7m30"><strong>eos7m30</strong></a></td><td>admet-ai-exact</td><td>ADMET properties prediction</td><td><a href="https://github.com/ersilia-os/eos7m30/blob/main/model/framework/columns/run_columns.csv">49 columns</a></td><td>ADMET AI is a framework for carrying out fast batch predictions for ADMET properties. It is based on ensemble of five Chemprop-RDKit models and has been trained on 41 tasks from the ADMET group in Therapeutics Data Commons (v0.4.1). Out of these 41 tasks, there are 31 classification tasks and 10 regression tasks. In addition to that output also contains 8 physicochemical properties, namely, molecular weight, logP, hydrogen bond acceptors, hydrogen bond doners, Lipinskis Rule of 5, QED, stereo centers, and topological polar surface area. eos7d58 contains an implementation of the model that also produces the percentile based on DrugBank approved drugs.</td><td><a href="https://github.com/ersilia-os/eos7m30">https://github.com/ersilia-os/eos7m30</a></td></tr><tr><td><a href="https://github.com/ersilia-os/eos9ei3"><strong>eos9ei3</strong></a></td><td>sa-score</td><td>Synthetic accessibility score</td><td><a href="https://github.com/ersilia-os/eos9ei3/blob/main/model/framework/columns/run_columns.csv">1 column</a></td><td>Estimation of synthetic accessibility score (SAScore) of drug-like molecules based on molecular complexity and fragment contributions. The fragment contributions are based on a 1M sample from PubChem and the molecular complexity is based on the presence/absence of non-standard structural features. It has been validated comparing the SAScore and the estimates of medicinal chemist experts for 40 molecules (r2 = 0.89). The SAScore has been contributed to the RDKit Package.</td><td></td></tr><tr><td><a href="https://github.com/ersilia-os/eos9yui"><strong>eos9yui</strong></a></td><td>natural-product-likeness</td><td>Natural product likeness score</td><td><a href="https://github.com/ersilia-os/eos9yui/blob/main/model/framework/columns/run_columns.csv">1 column</a></td><td>The model is a derivation of the natural product fingerprint (eos6tg8). In addition to generating specific natural product fingerprints, the activation value of the neuron that predicts if a molecule is a natural product or not can be used as a NP-likeness score. The method outperforms the NP_Score implemented in RDKit.</td><td></td></tr></tbody></table>

To explore the full list of models, visit the [Ersilia Model Hub browser](https://ersilia.io/model-hub).


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://ersilia.gitbook.io/ersilia-workshops/ai4dd-latam/sessions/monday-afternoon-breakout.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
