---
title: "Introduction to admetshiny"
author: "Xavier Clemente Garcia Cevallos"
date: "`r Sys.Date()`"
output:
  rmarkdown::html_vignette
vignette: >
  %\VignetteIndexEntry{Introduction to admetshiny}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r, include = FALSE}
knitr::opts_chunk$set(
  collapse = TRUE,
  comment = "#>",
  fig.width = 6, fig.height = 4
)
```

## Overview

**admetshiny** is an R package that provides an interactive Shiny application
and a toolbox of functions for the management, calculation, filtering,
visualization and exploratory analysis of molecular descriptors and ADMET
properties of small molecules. The application is organised in two
complementary modules:

1. **CDK & webchem** — retrieve canonical SMILES from PubChem, enter them
   manually or upload them as a CSV, then compute nine physicochemical
   descriptors locally with the Chemistry Development Kit (CDK).
2. **ADMET Master Manager** — upload any CSV or Excel (`.xlsx`) ADMET dataset
   and manually map its columns to the application's 20-field standard schema.
   Missing descriptors are back-filled from SMILES via CDK when available.

Both modules share the same drug-likeness filters (Lipinski, Veber, Ghose,
Egan, Muegge), the BOILED-Egg model, the P-gp substrate Random Forest
classifier and the 14-chart catalogue.

## Launching the application

The easiest way to use admetshiny is through its interactive application:

```r
admetshiny::run_app()
```

## Using the functions programmatically

The exported functions can also be used in plain R scripts.

### Drug-likeness filters

```r
library(admetshiny)

# Build a small toy dataset in the standard schema
d <- data.frame(
  MW = c(300, 650),
  LogP = c(2, 7),
  TPSA = c(40, 160),
  MR = c(70, 150),
  "#H-bond acceptors" = c(4, 12),
  "#H-bond donors" = c(2, 7),
  "#Rotatable bonds" = c(3, 14),
  "#Heavy atoms" = c(20, 80),
  "#Aromatic heavy atoms" = c(6, 9),
  check.names = FALSE
)

# Add the violation columns required by the filters
d <- computeViolationColumns(d)

# Apply Lipinski + Veber
filtered <- applyFilters(d, filters = c("Lipinski", "Veber"))
```

### CDK descriptors from SMILES

```r
library(admetshiny)

smiles <- c("CCO", "CC(=O)OC1=CC=CC=C1C(=O)O", "CN1C=NC2=C1C(=O)N(C(=O)N2C)C")

# Compute the 9 CDK descriptors (MW, ALogP, TPSA, HBD, HBA, RB, HA, AromHA, MR)
desc <- calcCDKDescriptors(smiles)

# Map to the standard schema and add #violations + ADMET properties
desc <- mapCDKDescriptors(desc)

# Apply all five drug-likeness filters
filtered <- applyFilters(desc,
                         filters = c("Lipinski", "Veber", "Ghose",
                                     "Egan", "Muegge"))
```

### Normalizing any external ADMET dataset

The `mapADMETColumns()` function replaces the former platform-specific
normalisation functions. It takes a raw data.frame, a user-specified named
mapping vector (column name -> standard field code), and an optional
`calculate_cdk` flag:

```r
d <- read.csv("my_admet.csv", check.names = FALSE)

# Map user columns to the standard schema. The codes are documented in
# ?mapADMETColumns. Missing descriptors are back-filled from SMILES via CDK.
mapping <- setNames(
  c("SMILES", "Name", "MW", "LogP", "TPSA"),
  c("CanonicalSMILES", "Compound", "MW", "iLOGP", "Topological PSA")
)
d <- mapADMETColumns(d, mapping, calculate_cdk = TRUE)
```

### BOILED-Egg plot

```r
# Requires LogP and TPSA columns. If a WLOGP column is present, the official
# WLOGP polygons are used; otherwise the ALogP-trained polygons.
plotBoiledEgg(filtered)
```

## Optional dependencies

Some features rely on suggested packages that are not installed automatically:

| Feature | Package |
|---|---|
| CDK descriptors | `rcdk` (requires Java JDK) |
| SMILES from PubChem | `webchem` |
| Radar plot | `fmsb` |
| t-SNE | `Rtsne`, `ggrepel` |
| UMAP | `uwot` |
| PCA labels | `ggrepel` |
| Tanimoto / AGNES | `rcdk`, `fingerprint`, `cluster` |
| Parallel coordinates | `GGally` |
| Excel upload / export | `openxlsx` |
| Colour palettes | `viridisLite` |

Install them with:

```r
install.packages(c("rcdk", "webchem", "fmsb", "Rtsne", "uwot", "ggrepel",
                   "fingerprint", "cluster", "GGally", "openxlsx",
                   "viridisLite"))
```

## References

- Lipinski, C. A., Lombardo, F., Dominy, B. W., & Feeney, P. J. (1997).
  *Advanced Drug Delivery Reviews*, 23(1-3), 3-25.
- Ghose, A. K., Viswanadhan, V. N., & Wendoloski, J. J. (1999). *J.
  Combinatorial Chemistry*, 1(1), 55-68.
- Veber, D. F., et al. (2002). *J. Medicinal Chemistry*, 45(12), 2615-2623.
- Egan, W. J., Merz, K. M., & Baldwin, J. J. (2000). *J. Medicinal Chemistry*,
  43(21), 3867-3877.
- Muegge, I., Heald, S. L., & Brittelli, D. (2001). *J. Medicinal Chemistry*,
  44(12), 1841-1846.
- Daina, A., & Zoete, V. (2016). A boiled egg to predict gastrointestinal
  absorption and brain penetration. *ChemMedChem*, 11(11), 1117-1121.
