Partial Least Squares Structural Equation Modeling, commonly known as PLS-SEM, has become an increasingly popular analytical method in PhD research, particularly in management, social sciences, education, engineering, urban studies, information systems, marketing, and behavioural research. One of the most widely used software applications for conducting PLS-SEM is SmartPLS.
For a beginner, SmartPLS can initially appear complex because it involves constructs, indicators, measurement models, structural models, reliability, validity, bootstrapping, path coefficients, effect sizes, and predictive measures. However, once the basic logic is understood, PLS-SEM becomes a highly systematic approach for testing relationships among latent variables.
This guide explains the essential concepts of SmartPLS and PLS-SEM for PhD scholars who are beginning to use the method.
What Is PLS-SEM?
PLS-SEM is a variance-based structural equation modelling technique used to examine relationships among latent constructs.
A latent construct is a concept that cannot be measured directly. Examples include:
customer satisfaction,
perceived safety,
accessibility,
service quality,
trust,
behavioural intention,
sustainability perception,
and organisational commitment.
Because these concepts cannot be observed directly, researchers measure them through several questionnaire items or indicators.
For example, a construct called “Public Transport Satisfaction” may be measured through statements related to comfort, punctuality, cleanliness, safety, and convenience.
PLS-SEM allows researchers to evaluate whether these indicators adequately measure the construct and whether the constructs are related in the way the theoretical model predicts.
Why Is PLS-SEM Popular in PhD Research?
PLS-SEM is attractive to PhD scholars because it can analyse complex models containing multiple constructs and several relationships at the same time.
It is often used when the research objective focuses on:
prediction,
explaining variance,
theory development,
testing complex conceptual models,
examining mediation,
analysing moderation,
and evaluating hierarchical constructs.
PLS-SEM is also relatively flexible regarding data distribution compared with some traditional covariance-based SEM approaches.
However, flexibility does not mean that PLS-SEM should be chosen automatically. The analytical method should always match the research objective and theoretical framework.
What Is SmartPLS?
SmartPLS is software designed primarily for Partial Least Squares Structural Equation Modeling.
It provides a graphical interface where researchers can:
import data,
create latent constructs,
assign indicators,
draw paths,
estimate PLS models,
perform bootstrapping,
assess reliability and validity,
analyse mediation and moderation,
and evaluate predictive performance.
This visual approach makes SmartPLS relatively accessible to researchers who may not have advanced programming skills.
Understanding the Measurement Model
One of the first concepts a PhD researcher must understand is the measurement model.
The measurement model explains how indicators relate to their latent constructs.
For example, suppose a study contains a construct called “Infrastructure Quality” measured using four questionnaire items:
IQ1
IQ2
IQ3
IQ4
SmartPLS estimates how strongly each item contributes to the construct.
These relationships are often represented through outer loadings.
Higher loadings generally indicate that an indicator represents the intended construct more strongly.
However, researchers should not remove indicators automatically simply because a loading is lower than expected. Indicator deletion should be based on both statistical evidence and theoretical justification.
Reflective and Formative Measurement Models
PLS-SEM distinguishes between reflective and formative measurement.
In a reflective model, the latent construct is assumed to influence its indicators.
For example, if a person has a high level of satisfaction, several satisfaction-related questionnaire items are expected to reflect that underlying satisfaction.
In a formative model, the indicators collectively form the construct.
For example, socioeconomic status might be formed by variables such as income, education, and occupation.
This distinction is extremely important because reflective and formative models require different evaluation procedures.
PhD scholars should decide the measurement type based on theory, not simply on software defaults.
Reliability Assessment
For reflective constructs, researchers commonly evaluate internal consistency reliability.
Important measures include:
Cronbach’s alpha,
rho_A,
composite reliability.
These statistics indicate whether the indicators consistently measure the same underlying construct.
Very low reliability values may indicate poor measurement quality.
However, extremely high reliability can also suggest that indicators are redundant.
Researchers should therefore interpret reliability together with theoretical relevance.
Convergent Validity
Convergent validity examines whether the indicators of a construct share sufficient common variance.
One widely used measure is Average Variance Extracted, or AVE.
An AVE value around or above 0.50 is commonly considered acceptable in many applications because it suggests that the construct explains at least half of the variance in its indicators.
However, researchers should not rely on AVE alone.
Outer loadings and reliability should also be considered.
Discriminant Validity
Discriminant validity determines whether two constructs are sufficiently distinct from one another.
If two supposedly different constructs are almost identical empirically, the model may have a conceptual or measurement problem.
The HTMT ratio is commonly used in PLS-SEM to assess discriminant validity.
Lower HTMT values generally indicate clearer separation between constructs.
Researchers should report the chosen threshold and justify it using appropriate methodological literature.
Collinearity Assessment
Before interpreting structural relationships, researchers should examine collinearity.
Variance Inflation Factor, or VIF, is commonly used for this purpose.
High VIF values may indicate that predictor constructs or indicators are excessively correlated.
This can make path coefficients unstable and difficult to interpret.
A good PhD thesis should report VIF values and explain whether multicollinearity is a concern.
Understanding the Structural Model
Once the measurement model is shown to be acceptable, researchers can evaluate the structural model.
The structural model represents the hypothesised relationships between constructs.
For example:
Accessibility → Public Transport Preference
Service Quality → Satisfaction
Satisfaction → Behavioural Intention
Each arrow represents a research hypothesis.
SmartPLS estimates a path coefficient for each relationship.
The sign indicates direction, while the magnitude indicates the strength of the relationship.
Bootstrapping in SmartPLS
Bootstrapping is one of the most important procedures in PLS-SEM.
It is used to assess the statistical significance of path coefficients and other model estimates.
SmartPLS repeatedly resamples the dataset and generates estimates of:
standard errors,
t-values,
p-values,
confidence intervals.
A path with a statistically significant result may support the proposed hypothesis.
However, statistical significance should not be interpreted in isolation.
Researchers should also consider effect magnitude, theoretical relevance, confidence intervals, and practical meaning.
R-Squared
R² indicates how much variance in an endogenous construct is explained by its predictors.
For example, if R² for Public Transport Preference is 0.60, the predictor constructs explain 60% of the variance in that outcome within the model.
Higher R² values indicate stronger explanatory power, but what counts as strong depends on the research field.
Researchers should avoid using arbitrary labels without considering disciplinary context.
Effect Size
The f² effect size helps determine how much an individual predictor contributes to the explained variance of an endogenous construct.
Two predictors may both be statistically significant but have very different practical importance.
Effect size therefore complements the path coefficient and p-value.
A PhD thesis should ideally discuss both significance and substantive importance.
Predictive Relevance
PLS-SEM is often used in prediction-oriented research.
Measures such as Q² and other predictive assessment tools can help evaluate whether the model has predictive relevance.
Recent PLS-SEM practice increasingly encourages researchers to examine prediction rather than relying only on in-sample explanatory measures.
This is particularly useful when the goal is to predict behaviour, preferences, adoption, or performance.
Mediation Analysis
Mediation occurs when one construct explains how or why another construct influences an outcome.
For example:
Service Quality → Satisfaction → Loyalty
Here, satisfaction may mediate the relationship between service quality and loyalty.
SmartPLS allows researchers to examine:
direct effects,
indirect effects,
total effects.
Bootstrapping is commonly used to evaluate the significance of mediation effects.
Researchers should explain the theoretical rationale for mediation before conducting the analysis.
Moderation Analysis
Moderation occurs when the strength or direction of a relationship changes depending on another variable.
For example, the effect of service quality on satisfaction might differ by age, income, or travel frequency.
SmartPLS can estimate interaction effects to test moderation.
Moderation should not be added merely to make a model appear more sophisticated. It should be based on a clear theoretical argument.
Sample Size in PLS-SEM
Sample size is a major concern in doctoral research.
Older PLS-SEM studies sometimes used simple rules such as the “10-times rule.”
However, researchers should avoid relying only on simplistic rules.
More rigorous approaches may involve:
statistical power analysis,
model complexity,
expected effect sizes,
significance level,
number of predictors,
and desired statistical power.
The sample should be justified clearly in the thesis methodology.
Data Preparation Before SmartPLS
Good analysis begins with clean data.
Before importing data into SmartPLS, researchers should check:
missing values,
incorrect coding,
duplicate responses,
reverse-coded items,
outliers,
measurement scale consistency,
and variable names.
Poor data preparation can create misleading model results.
The researcher should maintain a clear record of all data-cleaning decisions.
Common Mistakes Beginners Make
Several mistakes frequently occur among new SmartPLS users.
These include:
choosing PLS-SEM without theoretical justification,
deleting indicators only to improve reliability,
reporting only significant paths,
ignoring discriminant validity,
ignoring VIF,
confusing measurement and structural models,
using arbitrary sample-size rules,
interpreting p-values without effect sizes,
and modifying the model repeatedly until significant results appear.
Such practices can weaken the credibility of a PhD thesis.
The model should be driven by theory rather than by the desire to obtain significant findings.
How to Report SmartPLS Results in a Thesis
A clear reporting sequence is useful.
First, describe the conceptual model and hypotheses.
Second, explain the measurement model.
Third, report reliability and validity.
Fourth, assess collinearity.
Fifth, present structural model results.
Sixth, report path coefficients, confidence intervals, significance levels, R², effect sizes, and predictive measures where relevant.
Finally, interpret the findings in relation to theory and previous literature.
Tables and figures should support the explanation rather than replace it.
SmartPLS Is a Tool, Not a Research Strategy
A critical point for PhD researchers is that SmartPLS does not decide whether a research model is theoretically meaningful.
Software only estimates the model that the researcher specifies.
If the constructs are poorly defined, the questionnaire is weak, or the hypotheses lack theoretical support, sophisticated output cannot solve the problem.
A strong PLS-SEM study begins with:
theory → conceptual framework → hypotheses → measurement → data collection → model assessment → interpretation.
Research Integrity in PLS-SEM
Researchers should never manipulate their data simply to achieve significant path coefficients.
Non-significant relationships are valid research results.
Similarly, indicators should not be removed solely because their deletion improves statistical values unless there is a defensible theoretical reason.
Transparent reporting is essential.
The thesis should explain:
which indicators were retained or removed,
why changes were made,
how reliability and validity were assessed,
and whether hypotheses were supported or not supported.
Conclusion
SmartPLS and PLS-SEM can be powerful tools for PhD research when used appropriately.
They allow researchers to test complex relationships among latent constructs, evaluate measurement quality, estimate structural relationships, investigate mediation and moderation, and assess explanatory and predictive performance.
However, successful PLS-SEM research requires more than learning which buttons to click in SmartPLS.
PhD scholars need to understand the theoretical foundation of their constructs, select the correct measurement model, justify their sample, evaluate reliability and validity, assess structural relationships, and interpret results responsibly.
For doctoral researchers and academic support organisations such as EduPub, the most important principle is that SmartPLS should be used as an analytical tool within a sound research methodology.
A good PLS-SEM model is not the one with the largest number of significant paths. It is the one that is theoretically justified, methodologically transparent, statistically appropriate, and interpreted with academic integrity.

