A system of equations asks for the values of the variables that make every equation true at the same time. For two linear equations in two variables, there are three possibilities:
Graphs of the lines
Number of solutions
Name
Intersect once
1
consistent and independent
Parallel and distinct
0
inconsistent
Same line
infinitely many
consistent and dependent
The same idea extends to larger linear systems: the solution set can be one point, no points, or infinitely many points.
Geometrically, each equation describes a set of points. In two variables, each linear equation is a line, so solving a system means finding where the lines overlap. In three variables, each linear equation is usually a plane, so solving a system means finding where the planes overlap.
For a system in three variables:
one solution means the planes meet at one point,
no solution means the planes never all meet in one common place,
infinitely many solutions usually means the planes overlap along a line or plane.
Each method isolates the variables while preserving the set of points that satisfies every equation.
A matrix is a rectangular array of numbers. The numbers inside the matrix are called entries or elements.
Matrices are useful because linear systems have a lot of repeated structure. In a system like
{2xβy+4z=73x+5yβz=1β
the variable names x,y,z and the plus signs do not change much. The important changing information is the coefficients and constants. A matrix strips the system down to that information so the algebra becomes cleaner.
For example,
A=[20ββ13β45β]
has 2 rows and 3 columns, so its size is 2Γ3. The entry in row i and column j is often written aijβ. In the matrix above,
For a linear system, the coefficient matrix stores the coefficients of the variables. The augmented matrix also includes the constants on the right side.
The order of the columns matters. If the columns are arranged as x,y,z, then every row must follow that same order. Missing variables get coefficient 0. For example, 3x+z=5 becomes
Gaussian elimination is a systematic version of elimination. The main idea is to use one equation to remove a variable from the equations below it, then repeat with the next variable.
The reason this works is that each row in an augmented matrix is just one equation. If you replace an equation with a combination of equations that has the same information, the solution set stays the same.
For example, suppose a solution satisfies both
E1β:x+y=5
and
E2β:2xβy=4.
Then that same solution must also satisfy
E2ββ2E1β:(2xβy)β2(x+y)=4β2(5).
This new equation is not random. It is built from equations the solution already satisfies, so it does not throw away any valid solutions. Gaussian elimination keeps doing this kind of replacement until the system is easy to read.
The goal is usually to create row echelon form, which has a staircase of zeros:
ββ00βββ0ββββββββββ,
where the leading nonzero entries move down and to the right. Then use back substitution.
Why is this shape useful? Because the bottom equation has only one variable, the row above it has two variables, and the row above that has three variables. So you solve from the bottom upward.
A pivot is the entry you use to eliminate the numbers below it. In a typical 3Γ3 system, the first pivot is in the x-column. You use it to turn the entries below it into zeros:
It is often nice to make pivots equal to 1, but it is not required. The only thing a pivot cannot be is 0. If the entry where you want a pivot is 0, swap rows if possible.
Matrix operations are designed to preserve the rectangular structure of the data. The rules may feel more restrictive than ordinary arithmetic, but each restriction is there because the entries have positions. You can only combine entries when their positions match, and matrix multiplication has to respect row-column relationships.
You can add or subtract matrices only when they have the same size:
This is because addition is done position-by-position. The top-left entries combine, the top-right entries combine, and so on. If the matrices have different sizes, some entries do not have partners.
To multiply a matrix by a scalar, multiply every entry by that scalar:
Scalar multiplication stretches every entry by the same factor. If the matrix represents data, every data value is being scaled. If the matrix represents equations, multiplying a row by a scalar is the same idea as multiplying both sides of an equation by that scalar.
The product AB is defined only when the number of columns of A equals the number of rows of B.
Matrix multiplication is not entry-by-entry multiplication. It is built around dot products. The rows of the first matrix interact with the columns of the second matrix.
One reason this rule matters is that matrices often represent transformations or systems of linear combinations. Multiplying matrices combines those actions. If A changes one vector and B changes another, then AB represents doing one action after the other. Order matters, which is why AB and BA can be different.
If A is mΓn and B is nΓp, then AB is mΓp.
Each entry of AB is found by taking a row-column dot product:
A square matrix has the same number of rows and columns. The identity matrix is the matrix version of the number 1.
Multiplying by the identity matrix leaves a compatible matrix unchanged. An inverse matrix uses this property to βundoβ multiplication by another matrix.
For 2Γ2 matrices,
I2β=[10β01β].
For any compatible square matrix A,
AI=IA=A.
If a square matrix A has a matrix Aβ1 such that
AAβ1=Aβ1A=I,
then Aβ1 is the inverse of A. A matrix with an inverse is invertible or nonsingular. A matrix without an inverse is singular.
Think of Aβ1 as the operation that reverses A. If multiplying by A mixes the variables together, multiplying by Aβ1 unmixes them. This is why inverses can solve systems: the coefficient matrix mixes the variables into the constants, and the inverse recovers the original variables.
Not every square matrix can be undone. Some matrices collapse information. For example, if two equations are really multiples of the same equation, they do not contain enough independent information to recover a unique solution.
The determinant is a number calculated from a square matrix. A nonzero determinant means its rows and columns are linearly independent, so the matrix is invertible.
For a 2Γ2 coefficient matrix
[acβbdβ],
the two rows correspond to the coefficient patterns in two equations:
ax+by=constant,cx+dy=constant.
If the two rows point in genuinely different directions, the equations usually give two independent pieces of information and the system has one solution. If one row is a multiple of the other, the equations are parallel or identical, so the system either has no solution or infinitely many solutions.
The determinant detects this. For a 2Γ2 matrix,
det[acβbdβ]=adβbc.
If adβbc=0, the rows or columns are dependent in the sense that one direction has collapsed into another. The matrix is singular and has no inverse. If adβbcξ =0, the matrix is invertible.
Example. Compare the determinants:
A=[13β26β],B=[13β25β].
For A,
det(A)=1(6)β2(3)=0.
The second row is 3 times the first row, so the two rows do not give independent information.
For B,
det(B)=1(5)β2(3)=β1.
This is nonzero, so B is invertible and a system with coefficient matrix B has exactly one solution.
For a 3Γ3 matrix, one useful expansion is along the first row. Each entry in the first row gets multiplied by the determinant of the 2Γ2 matrix left behind after deleting that entryβs row and column:
If det(A)ξ =0, then A is invertible and the system AX=B has exactly one solution.
The determinant also has a geometric meaning. In two dimensions, the absolute value of the determinant gives the area scale factor of the matrix transformation. If the determinant is 0, area gets flattened to zero, meaning the transformation collapses the plane onto a line or point. That collapse is exactly why the matrix cannot be undone.
Cramerβs Rule is a determinant-based way to solve a system. It is not usually the fastest method for large systems, but it is useful because it shows how determinants encode the solution.
The denominator determinant D measures whether the coefficient matrix is invertible. The numerator determinants Dxβ and Dyβ replace one coefficient column at a time with the constants. This isolates how much of the solution belongs to each variable.
A nonlinear system has at least one equation that is not linear. The solutions are still points that satisfy every equation at once, but the graphs may intersect in more than one point.
Linear systems are predictable: two lines can meet once, never meet, or be the same line. Nonlinear systems are more flexible because curves can bend back and meet each other multiple times. A line and a circle can intersect twice, once, or not at all. Two circles can also intersect twice, once, or not at all.
The goal is still the same: find all ordered pairs that satisfy every equation. The difference is that the algebra often produces quadratics or higher-degree equations, so there may be multiple solutions.
Example. Solve
{x2+y2=25y=x+1β
Substitute y=x+1 into the circle equation:
x2+(x+1)2=25.
Then
2x2+2x+1=25,
so
2x2+2xβ24=0.
Divide by 2:
x2+xβ12=0.
Factor:
(x+4)(xβ3)=0.
Thus x=β4 or x=3. Since y=x+1, the solutions are
A system of inequalities asks for the region that satisfies every inequality at the same time.
Equations usually describe boundaries: lines, circles, parabolas, and so on. Inequalities describe regions on one side of those boundaries. A system of inequalities asks where all the shaded regions overlap.
For example, y>2x+1 means all points above the line y=2x+1. The line itself is not included because the inequality is strict. In contrast, yβ₯2x+1 includes the line.
Example. Describe the solution region:
{yβ₯2xβ1y<βx+5β
The first boundary is the solid line y=2xβ1. Shade above it.
The second boundary is the dashed line y=βx+5. Shade below it.
Find the original point (x,y) that maps to (7,5). Then find the image of the line y=2x+1 under this transformation.
We need
[21ββ11β][xyβ]=[75β].
This gives
{2xβy=7x+y=5β
Add the equations:
3x=12,
so x=4. Then y=1. The original point is
(4,1)β.
Now let points on the original line be written as
(x,y)=(x,2x+1).
The transformed coordinates satisfy
xβ²=2xβy=2xβ(2x+1)=β1,
and
yβ²=x+y=x+(2x+1)=3x+1.
As x varies, yβ² varies freely, but xβ² is always β1. Therefore the image of the line is
xβ²=β1β.
A small economy has two sectors: food and tools. Producing one unit of food requires 0.20 units of food and 0.10 units of tools. Producing one unit of tools requires 0.30 units of food and 0.20 units of tools. External demand is 110 units of food and 80 units of tools. Let F and T be the total production levels. Set up and solve the matrix equation for F and T.
The total production must satisfy internal demand plus external demand.
Food production:
F=0.20F+0.30T+110.
Tools production:
T=0.10F+0.20T+80.
Move internal demand terms to the left:
0.80Fβ0.30T=110,
and
β0.10F+0.80T=80.
In matrix form:
[0.80β0.10ββ0.300.80β][FTβ]=[11080β].
Clear decimals by multiplying both equations by 10: