Introduction

Mechanical tension is one of the main mechanisms inducing muscle hypertrophy by leading to signal transduction and increasing muscle protein synthesis (MPS) (Olsen et al., 2019; Wackerhage et al., 2018). In order to generate an optimal stimulus, several variables can be manipulated, being training volume, time under tension (TUT), frequency, load (generally expressed as percentages of the 1-repetition maximum) or proximity to muscle failure which is the most widely used (Bird et al., 2005). Proximity to failure is essential to achieve an optimal stimulus for muscle hypertrophy, regardless of the repetition range used, due to an increase in the recruitment of motor units (MUs) and their fatigue (Dankel et al., 2017; Morton et al., 2019). This proximity to failure could be managed by increasing the number of repetitions or by increasing the TUT within the same number of repetitions (Wilk et al., 2020). However, as long as the level of effort is high, training volume seems to be the most important variable (Schoenfeld et al., 2017). Therefore, when muscle hypertrophy is the main goal, training volume can be quantified as the number of sets per muscle group which are close to failure (Baz-Valle et al., 2018), namely, “hard sets”.

From an acute physiological standpoint, there is evidence suggesting a dose-response relationship between training volume and phosphorylation of proteins related to MPS (Gerasimos Terzis et al., 2010), or directly an increase in MPS (Burd et al., 2010). Interestingly, a recent study found greater responses in ribosomal biogenesis after an intervention of moderate training volume vs. low training volume (Hammarström et al., 2020). Ribosomal biogenesis, understood as ribosomal capacity, has been linked to long-term muscle mass gains along with ribosomal efficiency, since they are important physiological adaptations (Figueiredo, 2019). Moreover, this dose-response relationship has also been verified in longitudinal studies as a chronic response (Schoenfeld et al., 2019a), with a systematic review and meta-analysis (Schoenfeld et al., 2017) confirming these findings, with favorable results when performing over nine weekly sets per muscle group.

Over the last few years, training volume for muscle hypertrophy has received a lot of attention (Aube et al., 2020; Brigatto et al., 2019; Heaselgrave et al., 2019). Some studies support the dose-response hypothesis (Brigatto et al., 2019), while others propose an inverted “U” relationship between training volume and muscle mass gains (Heaselgrave et al., 2019). It is against this apparently contradictory background that this review intends to compare exclusively the response to moderate vs. high training volumes, in studies as homogeneous as possible, which include young trained men, and in which direct muscular hypertrophy measurements were taken. We hypothesized that the dose-response relationship would be minimal when comparing moderate and high training volumes.

Methods

Study design

A literature search of 3 databases was conducted in January, 2021. The following databases were searched: PubMed, Scopus and Cochrane Library. Databases were searched from inception up to January 2021, with no language limitation. Citations from scientific conferences were excluded.

Search strategy

The literature search was conducted in accordance with the Preferred Reporting Items for Systematic Reviews and Meta-analyses (PRISMA) guidelines. In each database, the title, abstract, and keywords search fields were searched. The following keywords, combined with Boolean operators (AND/OR) were used: “resistance training” AND Muscles AND hypertrophy OR “muscle thickness” AND volume. “Muscles” and “hypertrophy” were MeSH terms. No additional filters or search limitations were used. After conducting the initial search, the reference lists of articles retrieved were then screened for any additional articles which had relevance to the topic.

Eligibility criteria

Studies were eligible for further analysis if the following inclusion criteria were met: a) studies were randomized controlled trials comparing different groups with a different number of sets explicitly reported, with the same load assignment (%1-repetition maximum or XRM) and without the use of external implements (i.e., pressure cuffs, hypoxic chamber, etc.), b) interventions lasted at least six weeks, c) participants had a minimum of one year of resistance training experience, d) participants’ age ranged from 18 to 35 years, e) studies reported direct measurements of muscle thickness and/or the cross-sectional area, f) studies were published in peer-review journals.

Two independent observers reviewed the studies and then individually decided whether inclusion was appropriate. In the event of disagreement, a third observer was consulted. A flow chart of the search strategy and study selection is shown in Figure 1.

Figure 1

Flow diagram of the literature search.

https://jhk.termedia.pl/f/fulltexts/158681/j_hukin-2022-0017_fig_001_min.jpg

Study quality

Oxford’s level of evidence (OCEBM Levels of Evidence Working Group et al., 2011) and the Physiotherapy Evidence Database (PEDro) scale (de Morton, 2009) were used by two independent observers to assess the methodological quality of the studies included in the systematic review. Oxford’s level of evidence ranges from 1a to 5, with 1a being systematic reviews of high-quality randomized controlled trials and 5 being expert opinions. The PEDro scale consists of 11 different items related to scientific rigor. Given that the assessors are rarely blinded, and that it is hard to blind participants and investigators in supervised exercise interventions, items 5–7, which are specific to blinding, were removed from the scale (Baz-Valle et al., 2018). With the removal of these items, the maximum result on the modified PEDro 8-point scale was 7 (the first item was not included in the total score) and the lowest, 0. Zero points were awarded to a study that failed to satisfy any of the included items, and 7 pointed to a study that satisfied all the included items.

Included studies for qualitative and quantitative synthesis

In the present systematic review, seven studies met the inclusion criteria, in which low volume groups were included (<12 sets per week) descriptively in the qualitative synthesis, along with moderate (12-20 sets) and high training volume (>20 sets) groups. The main goal of this systematic review with meta-analysis was to compare moderate training volume vs. high training volume. Thus, six studies were included in the meta-analysis: those which included participants performing more than 12 weekly sets per muscle group, in order to compare moderate volumes (12-20 sets) vs. high volumes (>20 sets). Low volume groups (<12 sets) were excluded from the quantitative analysis.

Statistical analysis

Training interventions were classified as “high volume” (HV) if they included more than 20 weekly sets per muscle group, and as “moderate volume” (MV) otherwise. Groups in MV performed between 12 and 20 sets per muscle group in the included studies. Standardized mean difference (SMD) with 95% confidence intervals (CIs) between MV and HV regimens were calculated with RevMan 5.4 for macOS using the random effects model. Mean and SDs for the outcome measures were directly obtained from the original studies. The significance for an overall effect was set at p < 0.05. Heterogeneity of the analyzed studies was assessed using an I-squared test, setting the significance level at p < 0.01. The effects of each regimen (i.e., MV or HV interventions) were qualitatively assessed using the following threshold values for the SMD: 0.25, trivial; 0.25–0.50, small; 0.50–1.0, moderate; and >1.0, large. In some studies, more than one analysis was carried out because they included several groups performing HV (+20 weekly sets) (Brigatto et al., 2019), and/or performed different measurements to explore muscle changes in the case of quadriceps femoris (Aube et al., 2020). Three different analysis (one per muscle group) were performed to compare the effects of MV vs. HV in the different measurements. The first analysis explored effects of MV and HV in the quadriceps femoris muscle (including measurements of vastus lateralis, rectus femoris, and anterior thigh), while the second and the third analysis explored these effects in biceps brachii and triceps brachii muscle, respectively.

Results

Study selection

The search strategy yielded 2083 studies as presented in Figure 1. After removing 197 duplicates, and 1875 studies in the screening, 11 studies were determined to be potentially relevant to the topic based on the information contained in the abstract, from which only seven studies met the inclusion criteria. Excluded studies had at least one of the following characteristics: a) participants did not have enough training experience or had left their training programs long time ago and/or, b) there was only one training group, c) training sets were the same in each group, or d) the study was retracted from the journal (Figure 1). One study (Radaelli et al., 2015) did not meet the inclusion criteria because participants had insufficient specific training experience. Despite this, after analyzing the study (Radaelli et al., 2015) we realized that participants were trained in calisthenics and lifted their body weight performing 5RM in the bench press exercise, which suggested that they had a sufficient training level. Taking all data into account, all researchers agreed to include this study into the present systematic review with meta-analysis.

Finally, a total of seven studies which comprised 19 intervention groups, were included (Table 2). In the meta-analysis a total of six studies and 14 intervention groups were included. In all studies direct measurements were taken with ultrasounds (muscle thickness) and results were divided by measurements.

Level of evidence and quality of the studies

According to the Oxford’s level of evidence, four of the included studies had an evidence level 1b (high quality randomized controlled trials). The three remaining studies had a level of evidence 2b due to the following reason: less than 85% of participants completed the protocol. Scores from the PEDro scale were on average 4.7 ± 1.1, and ranged from 3 to 6 (Table 1). Quadriceps femoris qualitative analysis

In two out of the seven measurements (Brigatto et al., 2019; Schoenfeld et al., 2019a), significant differences between groups were observed, favoring high training volume for quadriceps hypertrophy vs. the low training volume group, and only in one study (Brigatto et al., 2019) significant differences between high and medium training groups were observed. In these studies, a larger effect size was observed, favoring the high training volume group. In the four remaining measurements (Amirthalingam et al., 2017; Aube et al., 2020), no significant differences were observed between groups, and the effect size did not favor any of the groups. Regarding the improvement percentage, in four out of the seven measurements (Brigatto et al., 2019; Schoenfeld et al., 2019a; Aube et al., 2020), larger gains were observed in the HV group (13.3, 9.4, 12.5 and 13.7% respectively) and, in the other three (Amirthalingam et al., 2017; Aube et al., 2020), in the MV groups (4.9, 6.9, 7.5% respectively).

Interestingly, Scarpelli et al. (2020) reported significant differences between groups in the increase in the cross-sectional area favoring the group that performed an individualized training volume. Regarding individual responses, ten participants (62.5% of the sample) had better responses when individualizing their training, two participants had a better response when not individualizing (12.5% of the sample), and four participants had a similar response (25% of the sample).

Quadriceps femoris quantitative analysis

The results classified between moderate and high volume are reported in Figure 2a. There were no significant effects for volume (p = 0.19); the effect size was -0.2 (CI: -0.49, 0.10), favoring high training volume. I2 = 0 value represents a high degree of homogeneity.

Figure 2

Forest plot of the comparison between MV and HV for quadriceps femoris measurements (a). Schoenfeld et al. (2019a) – measurements of rectus femoris for MV and HV groups. Schoenfeld et al. (2019b) – measurements of vastus lateralis for MV and HV groups. Brigatto et al. (2019a) – comparison between the MV group and HV1 group vastus lateralis measurements. Brigatto et al. (2019b) – comparison between the MV group and HV2 group vastus lateralis measurements. Aube et al. (2020a) – represents anterior thigh medial muscle thickness measurements. Aube et al. (2020b) – represents anterior thigh distal muscle thickness measurements. Aube et al. (2020c) – represents the sum of both anterior thigh muscle thickness measurements (medial and distal). Forest plot of the comparison between MV and HV for biceps brachii measurements (b). Brigatto et al. (2019a) – comparison between the MV group and HV1 group biceps brachii measurements. Brigatto et al. (2019b) – comparison between the MV group and HV2 group biceps brachii measurements. Forest plot of the comparison between MV and HV for triceps brachii measurements (c). Brigatto et al. (2019a) – comparison between the MV group and HV1 group triceps brachii measurements. Brigatto et al. (2019b) – comparison between the MV group and HV2 group triceps brachii measurements.

https://jhk.termedia.pl/f/fulltexts/158681/j_hukin-2022-0017_fig_002_min.jpg

Biceps brachii qualitative analysis

In two out of five studies (Radaelli et al., 2015; Schoenfeld et al., 2019a), significant differences were observed between groups, favoring high training volume. In one of them, significant differences were observed between HV and LV (Schoenfeld et al., 2019a), and in another study significant differences were observed between HV and MV, and between HV and LV (Radaelli et al., 2015). Among the remaining studies, a larger effect size between groups was observed in one of them, favoring the HV group (Brigatto et al., 2019); in another one, a larger effect size favoring the MV (Heaselgrave et al., 2019); and, in the last one, no significant differences were observed (Amirthalingam et al., 2017). Regarding the improvement percentage, larger gains in the HV group (17.5, 3, 6.9%) were observed in three out of five studies (Brigatto et al., 2019; Radaelli et al., 2015; Schoenfeld et al., 2019a), respectively, and in the MV group (7.2 and 8.5%) in the remaining two (Amirthalingam et al., 2017; Heaselgrave et al., 2019), respectively.

Biceps brachii quantitative analysis

Results classified between moderate and high volume are reported in Figure 2b. There were no significant effects for volume (p = 0.59). The effect size was -0,1 (CI: -0.46, 0.26), favoring high training volume. The I2 = 14 value represents a high degree of homogeneity.

Triceps brachii qualitative analysis

In two out of four studies (Brigatto et al., 2019; Radaelli et al., 2015), significant differences between groups were observed, favoring HV. In one of them (Brigatto et al., 2019), those differences were observed versus MV and, in another, versus both LV and MV (Radaelli et al., 2015). Among those studies in which no significant differences were observed between groups, in one of them a larger effect size was observed in HV compared to LV and MV (Schoenfeld et al., 2019a). Regarding the improvement percentage, a clear dose-response tendency in training volume and muscle mass gains was observed.

Triceps brachii quantitative analysis

Results classified between moderate and high volume are reported in Figure 2c. There were significant effects favoring high volume (p = 0.01); the effect size was -0.5 (CI: -0.88, 0.11), favoring high training volume. The I2 = 0 value represents a high degree of homogeneity.

Table 1

Physiotherapy Evidence Database (PEDro) ratings and Oxford evidence levels of the included studies

12345678TOTALEvidence level
Amirthalingam et al. (2017)Yes111101161b
Aube et al. (2020)Yes101001142b
Brigatto et al. (2019)Yes101101151b
Heaselgrave et al. (2019)Yes101101151b
Radaelli et al. (2015)Yes101111161b
Scarpelli et al. (2020)Yes101001142b
Schoenfeld et al. (2019)Yes100001132b
Total4.714

[i] Items in the PEDro scale: 1 = eligibility criteria were specified; 2 = subjects were randomly allocated to groups; 3 = allocation was concealed; 4 = the groups were similar at baseline regarding the most important prognostic indicators; 5 = measures of 1 key outcome were obtained from 85% of subjects initially allocated to groups; 6 = all subjects for whom outcome measures were available received the treatment or control condition as allocated or, where this was not the case, data for at least 1 key outcome were analyzed by “intention to treat”; 7 = the results of between-group statistical comparisons are reported for at least 1 key outcome; 8 = the study provides both point measures and measures of variability for at least 1 key outcome.