Dataset information
Available languages
English
Dataset description
The included files allow for the reconstruction of the data reported in: Notterpek, I., Craig, O.E., Garberi, P., Lucquin, A., Théry-Parisot, I., Abiven, S., 2025. BPChAr—a Benzene Polycarboxylic Acid database to describe the molecular characteristics of laboratory-produced charcoal: Implications for soil science and archaeology. PLoS One 20, e0321584. https://doi.org/10.1371/journal.pone.0321584 File name Description S1_Table.xlsx BPChAr database Sheet 1: Useful Information Table of Contents and useful information. Sheet 2: Bibliography Bibliography of publications screened and selected for inclusion in the database. Sheet 3: BPChAr database BPChAr database. Sheet 4: Treated_Data Pre-treated data for seamless application of the R code found in the Supporting Information (S1 Code). S2_Table.xlsx Additional information and results of statistical tests Sheet 1: Qualitative variables. Number of entries according to the qualitative variables here studied. Sheet 2: Pyrolysis temperature. Kruskal-Wallis rank sum test and Dunn’s post-hoc test results for all quantitative BPCA outputs as a function of pyrolysis temperature, categorised by low (≤ 300 °C, n = 68), mid (350 ≤ x < 700 °C, n = 139), and high (≥ 700 °C, n = 29) temperature chars. The degrees of freedom for each Dunn’s test is 2. Temperature categories with the same letter are not statistically distinct (p < 0.05). Sheet 3: Precursor feedstock. Results of two-way ANOVA with the Benjamini-Hochberg Procedure on all quantitative variables for low, mid, and high temperature chars among the precursor feedstock categories of hardwoods, softwoods, and grasses. Sheet 4: Oxygen availability during pyrolysis. Results of two-way ANOVA with the Benjamini-Hochberg Procedure on all quantitative variables for low, mid, and high temperature chars among the air composition categories “0,” “20.5,” “atmospheric,” and “atmospheric (restricted oxygen).” Sheet 5: Chromatographic separation method. Summary of statistical tests to investigate the effect of chromatographic separation method (GC- or LC-BPCA) for all, low, mid, and high temperature chars. The statistical test used was automatically determined in each case by the R code according to the sample size and results of the variable for the Shapiro-Wilk and Bartlett tests of normality. Statistically significant p-values are indicated in bold. Sheet 6: Random Forest results. Summary statistics for random forest predictive models testing various numbers of principal components, trees, and mtry parameters with PCA method 1 (omission) and 2 (imputation) for the treatment of missing values. In each case, the model with highest accuracy is indicated in bold. S1 Code: Notterpek_et_al_BPChAr.R R code utilised for the analysis of the BPChAr database. All functions and statistical tests presented in this work can be reproduced utilising this code and the data in S1 Table Sheet 4. Supplementary codes for proper functioning are provided in S2–S4 Code. S2 Code: ANOVA_2.R Integrated code for ANOVA and related tests for non-normally distributed data. S3 Code: Correlations.R Integrated code for correlation functions (e.g., Spearman correlation). S4 Code: Fonction_Predictions.R Integrated code necessary for the random forest models. This project received funding from the European Union’s Horizon 2020 research and innovation programme under the Marie Skłodowska-Curie grant agreement No 956351. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.
European data infrastructure with broad catalog discovery, free evaluation access and production-grade API options.
190K+
indexed dataset pages
32
countries and EU institutions
Free API quota
for evaluation and prototypes
SLA
history and push on production APIs