The Unabridged Pentium 4. IA32 Processor Genealogy
Book information
Description
Title Page At-a-Glance Table of Contents Table of Contents List of Figures List of Tables Acknowledgements About This Book The IA32 Architecture Specification The Pentium® 4 Is the Sum of Its Ancestors The CD The MindShare Architecture Series Cautionary Note The Specification Is the Final Word Documentation Conventions Hexadecimal Notation Binary Notation Decimal Notation Bits Versus Bytes Notation Bit Fields (Logical Groups of Bits or Signals) Signal Names Visit Our Web Site We Want Your Feedback Part 1 Introduction 1 Overview of the Processor Role The IA32 Specification IA32 Processors IA32 Instructions vs. µops Processor = Instruction Fetch/Decode/Execute Engine Some Instructions Result in FSB Transactions Many Instructions Do Not Require FSB Transactions Instructions That Do Require FSB Transactions IO Read and Write IO Read Instruction IO Write Instruction Memory Data Read Memory Data Write Memory Instruction Read The Processor’s Role in Today’s Systems Processor Activities at Startup Processor Activities During Run-Time Load and Run Application Programs Application Program Calls the OS Handling External Hardware Interrupts Calling a Device Driver Handling Software Exceptions System Overview Pentium® 4 Processor Memory Control Hub (MCH) IO Control Hub (ICH) Super IO (SIO) Chip DDR RAM IDE RAID Controller USB 2.0 Controller Five PCI Card Slots Part 2 Single-/MultiTask OS Background 2 Single-Task OS and Application Operating System Overview Command Line Interpreter (CLI) Program Loader OS Services Direct IO Access Application Program Memory Usage Task Initiation, Execution and Termination 3 Definition of Multitasking Concept An Example-Timeslicing Another Example-Awaiting an Event Task Issues Call to OS for Disk Read OS Suspends Task OS Initiates Disk Read OS Makes Entry in Event Queue OS Starts or Resumes Another Task Disk-Generated Interrupt Causes Jump to OS Task Queue Checked OS Resumes Task 4 Multitasking Problems OS Protects Territorial Integrity Stay in Your Own Memory Area IO Port Anarchy Unauthorized Use of OS's Tools No Interrupts, Please! BIOS Calls Part 3 The 386 5 386 Real Mode Operation Special Note An Overview of the 386 Internal Architecture An Overview of the 386DX FSB Address Bus Selects Dword Byte Enables Select Location(s) in Dword Misaligned Transfers Affect Performance Alignment Is Important! The 386 Register Set Control Registers CR0 CR1 CR2 CR3 EFlags Register General Purpose Registers (GPRs) Introduction EAX, EBX, ECX and EDX Registers EBP Register Index Registers Segment Registers Real Mode Usage Protected Mode Usage DS, ES, FS and GS CS SS Extended Instruction Pointer (EIP) Register Task Register What Is a TSS? The Purpose of the Task Register GDTR and LDTR IDTR Hardware Interrupts Software Exceptions and Interrupts IDTR Points To the Interrupt Table The Debug Registers Test Registers 386 Power-Up State Initial Memory Reads IO Port Addressing Memory Addressing General Accessing the Code Segment Accessing the Stack Segment Introduction Pushing Data Onto the Stack Popping Data From the Stack Processor Stack Usage Accessing the DS Data Segment Accessing the ES/FS/GS Data Segments An Example Accessing Extended Memory in Real Mode Big Real Mode Real Mode Instructions and Registers Registers Accessible in Real Mode Registers Inaccessible in Real Mode Instructions Usable in Real Mode Instructions Unusable in Real Mode Real Mode Interrupt/Exception Handling Protection in Real Mode 6 Protected Mode Introduction General Memory Protection Segmentation Virtual Memory Paging IO Protection Privilege Levels Virtual 8086 Mode Task Switching Interrupt Handling Real Mode Interrupt Handling Protected Mode Interrupt Handling 7 Intro to Segmentation in Protected Mode Special Note Real Mode Limitations Segment Descriptor Describes a Memory Area in Detail Segment Register-Selects Descriptor Table and Entry Introduction to the Descriptor Tables Segment Descriptors Reside in Memory Global Descriptor Table (GDT) GDT Description Setting the GDT Base Address and Size Local Descriptor Tables (LDTs) General Segment Descriptor Format Granularity Bit Segment Base Address Field Segment Size Field Default/Big Bit In a Code Segment, It’s the Descriptor’s “Default” Bit In a Stack Segment, It’s the Descriptor’s “Big” Bit Segment Type Field Introduction to the Type Field Non-System Segment Types Segment Present Bit Descriptor Privilege Level (DPL) Field System Bit Available Bit 8 Code Segments Selecting the Code Segment to Execute Code Segment Descriptor Format Accessing the Code Segment Privilege Checking General Some Definitions Definition of a Task Definition of a Procedure CPL Definition DPL Definition Conforming and Non-Conforming Code Segments RPL Definition Calling a Procedure in the Current Task Call Gate The Problem The Solution-Different Gateways Call Gate Example Execution Begins Call Gate Descriptor Read Call Gate Contains Code Segment Selector Code Segment Descriptor Read The Big Picture The Call Gate Privilege Check Privilege Check for a Call through a Call Gate Privilege Check for a Jump through a Call Gate Automatic Stack Switch 9 Data and Stack Segments A Note Regarding Stack Segments The Data Segments Selecting and Accessing a Data Segment Data Segment Privilege Check Selecting and Accessing a Stack Segment Introduction Expand-Up Stack Expand-Down Stack The Problem Expand-Down Stack Description An Example Another Example Stack Segment Privilege Check 10 Creating a Task What Is a Task? Basics of Task Creation and Startup Load All or Part of the Task into Memory Create a TSS and a TSS Descriptor for the Task Trigger the Timeslice Timer Scheduler Causes a Task Switch Interrupt on Timer Expiration TSS Structure General IO Port Access Protection IO Protection in Real Mode Definition of IO Privilege Level (IOPL) IO Permission Check in Protected Mode IO Permission Check in VM86 Mode IO Permission Bit Map Offset Field Interrupt Redirection Bit Map OS-Specific Data Structures Debug Trap Bit (T) LDT Selector Field Segment Register Fields General Register Fields Extended Stack Pointer (ESP) Register Field Extended Flags (EFlags) Register Field Extended Instruction Pointer (EIP) Register Field Control Register 3 (CR3) Field Privilege Level 0 - 2 Stack Definition Fields Link Field (to Old TSS Selector) TSS Descriptor How the OS Starts a Task What Happens When a Task Starts Use of the LTR and STR Instructions General The STR Instruction The LTR Instruction 11 Mechanics of a Task Switch Events that Initiate a Task Switch Switch Via a TSS Descriptor Task Gate Descriptor Task Gate Selected by a Far Call/Jump Task Gate Selected by a Hardware Interrupt or a Soft ware Exception Task Gate Selected by an INT Instruction Task Switch Details Switch Due To an Interrupt or Exception Switch as a Result of a Far Call Switch as the Result of a Far Jump Switch Due to a BOUND/INT/INTO/INT3 Instruction Switch Due to Execution of an IRET Linked Tasks Linkage Modification The Busy Bit Address Mapping The Linear vs. the Physical Memory Address The GDT Purpose and Location The LDT Purpose and Location Paging-Related Issues Background Each Task Can Have Different Linear-to-Physical Mapping TSS Mapping Must Remain the Same for All Tasks Placement of a TSS Within a Page(s) 12 386 Demand Mode Paging Problem-Loading Entire Task into Memory is Wasteful Solution-Load Part and Keep Remainder on Disk Load on Demand Track Usage Capabilities Required Problem-Running Two (or more) DOS Programs Solution-Redirect Memory Accesses to Separate Memory Areas Global Solution-Map Linear Address to Disk Address or to a Different Physical Memory Address The Paging Unit Is the Translator Linear Memory Space Is Divided into 220 4KB Pages Physical Memory Space Is Divided into 220 4KB Pages Mass Storage Space Is Divided into 4KB Pages The Paging Unit Uses Directories to Remap the Address Three Possible Page Lookup Methods First Method: Sequential Scan through a Large Table Second Method: Index into a Large Table Third Method: Index into a Selected Small Table IA32 Page Lookup Method Enabling Paging Page Directory and Page Tables Finding the Location of a Physical Page Find the Page Table First When the Target Page Table Is in Memory When the Target Page Table Isn’t in Memory Find the Page Using an Entry in a Page Table When the Target Page Is in Memory When the Target Page Isn’t in Memory Eliminating the Directory Lookup The 386 TLB TLB Maintenance The TLBs Are Cleared on a Task Switch or a Page Directory Change Updating a Single Page Table Entry Checking Page Access Permission The Privilege Check Segment Privilege Check Takes Precedence Over Page Check U/S Bit in Page Directory and Page Table Entry Is Checked Accesses with Special Privilege The Read/Write Check Page Faults Page Fault Causes Second Page Fault while in the Page Fault Handler A Page Fault During a Task Switch A Page Fault while Changing to a Different Stack Page Fault Error Code Usage of the Dirty and Accessed Bits Demand Mode Paging Evolution 13 The Flat Model Segments Complicate Things Paging Can Do It All Eliminating Segmentation The Privilege Check The Read/Write Check Each Task (including the OS) Has Its Own TSS Switch to an Application Task Switch to an OS Kernel Task 14 Interrupts and Exceptions Special Note General Hardware Interrupts Maskable Interrupt Requests Maskable Interrupt Servicing Automatic Actions Actions Performed by the Software Handler PC-Compatible Vector Assignment Non-Maskable Interrupt Requests Software-Generated Exceptions General Faults, Traps, and Aborts Instruction Restart Software Interrupt Instructions Interrupt/Exception Priority Real Mode Interrupt/Exception Handling Real Mode Interrupt Descriptor Table (IDT) Structure Real Mode Interrupt/Exception Handling Protected Mode Interrupt/Exception Handling General Protected Mode Interrupt Descriptor Table (IDT) Structure Handlers Can Only Be Entered From Program of Equal or Lesser Privilege Lower-Privilege Programs Could Call More Privileged Programs Gates Prevent Anarchy Interrupts/Exceptions Bypass the Gate Privilege Check Interrupt Gates General Actions Taken when an Interrupt Selects an Interrupt Gate Trap Gates Using a Procedure as an Interrupt/Exception Handler State Save Jump to the Handler Return to the Interrupted Program Returning to the Same Privilege Level Returning to a Different Privilege Level Using a Task as an Interrupt/Exception Handler Interrupt/Exception Handling in VM86 Mode Exception Error Codes The Resume Flag Prevents Multiple Debug Exceptions Special Case-Interrupts Disabled While Updating SS:ESP The Problem The Solution Detailed Description of the Software Exceptions Divide-by-Zero Exception (0) Processor Introduced In Exception Class Description Error Code Saved Instruction Pointer Processor State Debug Exception (1) Processor Introduced In Exception Class Description Error Code Saved Instruction Pointer Processor State NMI (2) Processor Introduced In Exception Class Error Code Saved Instruction Pointer Processor State Breakpoint Exception (3) Processor Introduced In Exception Class Description Error Code Saved Instruction Pointer Processor State Overflow Exception (4) Processor Introduced In Exception Class Description Error Code Saved Instruction Pointer Processor State Array Bounds Check Exception (5) Processor Introduced In Exception Class Description Error Code Saved Instruction Pointer Processor State Invalid OpCode Exception (6) Exception Class Description Error Code Saved Instruction Pointer Processor State Device Not Available Exception (7) Processor Introduced In Exception Class Description Error Code Saved Instruction Pointer Processor State Double Fault Exception (8) Processor Introduced In Exception Class Description Shutdown Mode Error Code Saved Instruction Pointer Processor State Coprocessor Segment Overrun Exception (9) Processor Introduced In Exception Class Description Error Code Saved Instruction Pointer Processor State Invalid TSS Exception (10) Processor Introduced In Exception Class Description Error Code Saved Instruction Pointer Processor State Segment Not Present Exception (11) Processor Introduced In Exception Class Description Error Code Saved Instruction Pointer Processor State Stack Exception (12) Processor Introduced In Exception Class Description Error Code Saved Instruction Pointer Processor State General Protection (GP) Exception (13) Processor Introduced In Exception Class Description Error Code Saved Instruction Pointer Processor State Page Fault Exception (14) Processor Introduced In Exception Class Description Error Code CR2 Saved Instruction Pointer Processor State The More Common Case Page Fault During a Task Switch Page Fault During a Stack Switch Vector 15 FPU Exception (16) Processor Introduced In Exception Class Description Handling of Masked Errors Handling of Unmasked Errors Error Code Saved Instruction Pointer Processor State Alignment Check Exception (17) Processor Introduced In Exception Class Description Implicit Privilege Level 0 Accesses Storing GDTR, LDTR, IDTR or TR FP/MMX/SSE/SSE2 Save and Restore Accesses MOVUPS and MOVUPD Accesses FSAVE and FRSTOR Accesses Error Code Saved Instruction Pointer Processor State Machine Check Exception (18) Processor Introduced In Exception Class Description Error Code Saved Instruction Pointer Processor State SIMD Floating-Point Exception (19) Processor Introduced In Exception Class Description Exception Error Code Saved Instruction Pointer Processor State 15 Virtual 8086 Mode A Special Note DOS Application-Portrait of an Anarchist Solution-Set a Watchdog on the DOS Application The Virtual Machine Monitor (VMM) Entering or Reentering VM86 Mode Task Creation, Startup and Suspension Create a TSS Each Task Gets a Timeslice Select DOS Task via a Far Call or a Far Jump An Interrupt or Exception Causes an Exit From VM86 Mode General An Interrupt or Exception Clears EFlags[VM] IRET Sets EFlags[VM] Again A Task Switch Causes an EFlags Update DOS Task's Memory Usage 1st MB Is DOS Memory Paging Provides Each DOS Task with Its Own Copy of the 1st MB The VMM Should Not Reside in the HMA Dealing with Segment Wraparound 8088/8086 Processor Post-8086 Processors Solutions Segment Register Interpretation in VM86 Mode Using the Address Size Override Prefix The Privilege Level of a VM86 Task Restricting IO Accesses The Problem IO-Mapped IO IO Permission in Protected Mode IO Permission in VM86 Mode Memory-Mapped IO Segregate Ports into Two Groups of Memory Pages Set Up Task’s Page Tables to Permit or Deny Access Handling Display Frame Buffer Updates IOPL-Sensitive Instructions The Problem-Instructions with Side Effects CLI (Clear Interrupt Enable) Instruction STI (Set Interrupt Enable) Instruction PUSHF (Push Flags) Instruction POPF (Pop Flags) Instruction INT nn (Software Interrupt) Instruction IRET (Interrupt Return) Instruction The Solution-IOPL Sensitive Instructions Interrupt/Exception Generation and Handling Introduction Normally, There’s Only One IDT VM86 Mode-There Are Two IDTs Which IDT Is Used? Processor Actions when a Hardware Interrupt Occurs in VM86 Mode Obtain the Vector from the Interrupt Controller The Vector Selects a Protected Mode IDT Entry Critical Values Are Stored on VM86 Task’s Level 0 Stack Jump to the Handler The Handler May Expect Values in the Data Segment Registers A Handler May Need to Know It Was Entered from a VM86 Task A Handler May Need to Return Values in the Data Segment Registers Exit the Handler and Return to the Interrupted VM86 Task Why the Data Segment Registers Were Cleared Execution of the IRET Instruction Processor Actions When an INT nn Is Executed in VM86 Mode When the IOPL < 3 When the IOPL = 3 Processor Actions when an Exception Occurs in VM86 Mode Execute the Protected Mode Handler or Pass Control to the VMM The VMM Chooses Its Actions Based on the Vector The VMM Passes the Ball to a Real Mode Handler Sometimes, the VMM Handles the Event General Attempt to Access a Forbidden IO Port Attempted Execution of a CLI Instruction Attempted Execution of the STI Instruction Attempted Execution of a PUSHF Instruction Attempted Execution of a POPF Instruction Attempted Execution of the INT nn Instruction Attempted Execution of an IRET Instruction Using a Separate Task as a Handler in VM86 Mode Registers Accessible in Real/VM86 Mode Instructions Usable in Real/VM86 Mode VM86 Mode Evolution 16 The Debug Registers The Debug Registers Part 4 486 17 Caching Overview Definition of a Load and a Store The Cache’s Purpose Without a Cache, Core Stalls Were Common An On-Die Cache Eliminates Many Core Stalls Introduction On a Cache Miss The Cache Line The Directory Entry Repeat Accesses to the Same Areas Result in Cache Hits The Write-Through Cache Introduction On a Load Miss On a Load Hit On a Store Miss On a Store Hit Additional Information on the WT Cache The Write Back Cache A Line Can Be in One of Four Possible States Before Storing to a Shared Line, Kill All Other Copies On a Store Miss, Perform an RWITM Additional Information on the WB Cache Snooping General Snooping and the WT Cache Introduction Snooping a Memory Read in a WT Cache Snooping a Memory Write in a WT Cache Snooping and the WB Cache Introduction Snooping a Memory Read in a WB Cache Snooping a Memory Write in a WB Cache The Overall Cache Architecture Introduction The Fully-Associative Cache Two-Way Set Associative Cache Four-Way Set Associative Cache Eight-Way Set Associative Cache Cache Real Estate Management The Lookup The Cache Initiates the Fetch of the New Line And Immediately Decides Where to Store It Example 1: Castout of a Modified Line Example 2: Castout of an E or S Line A Unified Cache Split Caches Non-Blocking Caches 18 486 Hardware Overview 486 Flavors An Overview of the 486 Internal Architecture An Overview of the 486 FSB Address/Data Bus Structure On a Cache Miss, an Entire Line Must Be Fetched 486 Implemented a Burst Line Fill Transaction Background Toggle Mode Transfer Order A20 Mask Accessing Extended Memory in Real Mode Segment Wraparound 486 Integrated the A20M# Gate On-Chip Cache Added General Cache Operation An Example On a Memory Read (i.e., a Load) Lookup On a Memory Write (i.e., a Store) Lookup 19 486 Software Enhancements FPU Added On-Die Introduction FPU-Related Register Set Changes The CR0 FPU Control Bits The FP Data Registers The FCW Register The FSW Register The FTW Register The Instruction Pointer Register The Data Pointer Register The FP Data Operand Format FP Error Reporting Precise Error Reporting Imprecise Error Reporting Why Deferred Error Reporting Is Used The WAIT/FWAIT Instruction The NE Bit DOS-Compatible FP Error Reporting FP Error Reporting Via Exception 16 Ignoring FP Errors Alignment Checking Feature Paging-Related Changes The Write Protect Feature Description Example Usage: Unix Copy-on-Write Strategy Directory/Table and Page Caching Page Directory Caching Page Table Caching Page Caching Caching-Related Changes to the Programming Environment CR4 Was Added in the Later Models of the 486 Test Registers Added Instruction Set Changes Exchange and Add (XADD) Compare and Exchange (CMPXCHG) Invalidate Cache (INVD) Write Back and Invalidate (WBINVD) Invalidate TLB Entry (INVLPG) Resume from System Management Mode (RSM) Byte Swap (BSWAP) New/Altered Exceptions Exception 9 Is Now Reserved Exception 17 (Alignment Check) Added System Management Mode (SMM) Part 5 Pentium® 20 Pentium® Hardware Overview Pentium® Flavors An Overview of the Pentium® Internal Architecture The First Superscalar IA32 Processor Brief Core Description An Overview of the Pentium® FSB Address/Data Bus Structure Address Bus Selects Qword Byte Enables Select Location(s) in Qword On a Cache Miss, an Entire Line Must Be Fetched The Burst Transaction Background Toggle Mode Transfer Order Burst Write Transaction The Caches Split Cache Structure Pentium® Code Cache Pentium® Data Cache Local APIC Added in the P54C Test Access Port (TAP) General Operational Description FRC Mode Soft Reset (INIT#) Hot Reset and 286 DOS Extender Programs Alternate (Fast) Hot Reset 286 DOS Extenders on Post-286 Processors 21 Pentium® Software Enhancements VM86 Extensions Introduction Efficient CLI/STI Instruction Handling Background CLI Handling STI Handling Efficient Handling of the INT Instruction Protected Mode Virtual Interrupts Debug Extension Time Stamp Counter Reading the TSC Writing to the TSC Restricting Access to the TSC Counter Wraparound RDTSC Is Not a Serializing Instruction 4MB Pages The Problem How To Set Up a 4MB Page Other PDEs Can Point to Page Tables The Address Translation Machine Check Architecture (MCA) Performance Monitoring Local APIC Register Set The Problem The Solution The Local APIC’s Register Set The Pentium® Local APIC’s Characteristics Detailed Description of the APIC Test Registers Relocated MSRs Added General Test Register 12 Instruction Set Changes Non-MMX Instructions CMPXCHG8B RDTSC RDMSR and WRMSR RDMSR WRMSR CPUID Instruction Description MMX Capability Introduction The Basic Problem MMX SIMD Solution Dealing with Unpacked Data Dealing with Math Underflows and Overflows Elimination of Conditional Branches Introduction Non-MMX Chroma-Key/Blue Screen Compositing Example MMX Chroma-Keying/Blue Screen Compositing Example Detecting MMX Capability Changes To the Programming Environment Handling a Task Switch MMX Instruction Set Syntax MMX Execution Unit New/Altered Exceptions Exception 13d Exception 14d Exception 18d Part 6 Intro to the P6 Core and FSB 22 P6 Road Map The P6 Processor Family The Klamath Core The Deschutes Core The Katmai Core 23 P6 Hardware Overview For More Detail Introduction The P6 Processor Core The FSB Interface Unit The Agent Types The Request Agent Types The Transaction Phases The Transaction Types The Backside Bus (BSB) Interface Unit The Unified L2 Cache The L1 Data Cache The L1 Code Cache The Processor Core The Local APIC Unit Part 7 Pentium® Pro Software Enhancements 24 Pentium® Pro Software Enhancements Paging Enhancements PAE-36 Mode The Problem The Solution: PAE-36 Mode Enabling PAE-36 Mode The Application Is Still Limited to a 4GB Virtual Address Space The OS Creates the Application’s Address Translation Tables CR3 Is Loaded with the Top Level Address Translation Table Pointer The Page Directory Pointer Table Lookup The Page Directory Lookup PDE Points to a Page Table PDE Points to a 2MB Physical Page The Page Table Lookup Windows OS PAE Support Linux PAE Support Global Pages Problem Global Page Feature APIC Enhancements MMX Not Implemented SMM Enhancement MTRRs Added Know the Characteristics of Your Target Introduction Why the Processor Must Know the Memory Type Earlier CPUs Required Chipset Memory Type Registers The Memory Type Registers Are Now Part of the CPU Architecture MTRRs Are Divided Into Four Categories MTRR Feature Determination MTRRDefType Register State of the MTRRs after Reset The Fixed Range MTRRs The Problem Enabling the Fixed Range MTRRs They Define the Memory Types Within the 1st MB of Memory Space The Variable-Range MTRRs Enabling the Variable-Range MTRR Register Pairs The Number of Variable-Range MTRR Register Pairs The Format of the Variable-Range MTRR Register Pairs The MTRRPhysBasen Register The MTRRPhysMaskn Register Variable-Range Register Pair Programming Examples The Memory Types Uncacheable (UC) Memory Write-Combining (WC) Memory Write-Through (WT) Memory Write-Protect (WP) Memory Write-Back (WB) Memory Rules as Defined by MTRRs Rules of Conduct Provided in Bus Transaction Paging Also Defines the Memory Type MTRRs Must Be the Same in an MP System MCA Enhanced MCA = Error Logging Capability The MCA Elements The Machine Check Exception The MCA Register Set The Global Registers Introduction The Global Count and Present Register The Global Status Register The Global Control Register The Composition of a Register Bank Overview The Bank Control Register The Bank Status Register General Error Valid Bit Overflow Bit Uncorrectable Error Bit Error Enabled Bit Miscellaneous Register Valid Bit Address Register Valid Bit Processor Context Corrupt Bit MCA Error Code and Model Specific Error Code Other Information The Bank Address Register The Bank Miscellaneous Register The Error Code The Error Code Fields Simple MCA Error Codes Compound MCA Error Codes FSB Error Interpretation MC Exception May or May Not Be Recoverable Machine Check and BINIT# Additional Error Logging Notes Error Buffering Capability Additional Information for Each Log Entry Initialization of the MCA Register Set The Performance Counters Purpose of the Performance Monitoring Facility Performance Monitoring Registers PerfEvtSel0 and PerfEvtSel1 MSRs PerfCtr0 and PerfCtr1 Accessing the Performance Monitoring Registers Accessing the PerfEvtSel MSRs Accessing the PerfCtr MSRs Accessing Using RDPMC Instruction Accessing Using RDMSR/WRMSR Instructions Event Types Starting and Stopping the Counters Starting the Counters Stopping the Counters Performance Monitoring Interrupt on Overflow MSRs Added Some Notes Test Control Register (TEST_CTL) ROB_CR_BKUPTMPDR6 MSR DebugCtl MSR General BPM and BP Pin Usage Enable Branch Trace Messaging The Branch, Exception, Interrupt Recording Facility General LastBranchFromIP and LastBranchToIP Register Pair LastExceptionFromIP and LastExceptionToIP Register Pair Single-Step on Branch, Exception, or Interrupt Instruction Set Changes MMX Not Implemented New Instructions Conditional Move (CMOV) Eliminates Branches Problem It Addresses Description Conditional FP Move (FCMOV) Eliminates Branches Problem Addressed Description FCOMI, FCOMIP, FUCOMI, and FUCOMIP RDPMC Problem Addressed Description UD2 The CPUID Instruction Enhanced New/Altered Exceptions 25 MicroCode Update Feature The Problem The Solution The Microcode Update Image Introduction The Microcode Update Header Matching the Image to a Processor CPUID Enhanced to Supply Update Signature Processor/Image Match Determination The Microcode Update Loader The Bare Bones Loader The Trigger Initiates the Upload Process After the Upload, the Signature Is Updated Authenticating the Image Additional Loader Requirements Possible Loader Enhancements Updates in a Multiprocessor System The Image Management BIOS The Purpose of the Image Management BIOS The BIOS Interface Detailed Function Call Description The Presence Detect Function Call The Write Microcode Update Data Function Call The Microcode Update Control Function Call The Read Microcode Update Data Function Call When Must the Image Upload Take Place? Determining if a New Update Supersedes a Previously- Loaded Update Effect of RESET# Or INIT# on a Previously-Loaded Update Part 8 Pentium® II 26 Pentium® II Hardware Overview The Pentium® Pro and Pentium® II: Same CPU, Different Package Dual-Independent Bus Architecture (DIBA) IOQ Depth Pentium® Pro/Pentium® II Differences One Product Yields Three Product Lines The Pentium® II/Xeon/Celeron Roadmap The Cartridge The Pentium® and Pentium® Pro Sockets The Problem The Pentium® II Cartridge The SEC Substrate: the Processor Side General Processor Core The SEC Substrate: the Non-Processor Side Cartridge Block Diagram The L2 Cache The Core General L1 Caches L1 Code Cache Characteristics L1 Data Cache Characteristics L1 and L2 Cache Error Protection 16-bit Code Optimization The Pentium® Pro Was Not Optimized Pentium® II Shadows the Data Segment Registers The FSB and BSB The FSB Protocol The Processor Core and Bus Frequencies The FSB Arbitration Scheme The Pentium® Pro Processor FSB Arbitration Pentium® II Processor FSB Arbitration The BSB and the L2 Cache The BSB Frequency The L2 Cache The Introduction of the Celeron Miscellaneous Hardware Stuff Pentium® II/Pentium® Pro Signal Differences Voltage Identification 27 Pentium® II Power Management Features The Pentium® Pro’s Power Conservation Modes The Pentium® II’s Power Conservation Modes The Normal State The AutoHalt Power Down State Description The Chipset’s Response to the Halt Message The Stop Grant State The Halt/Grant Snoop State The Sleep State The Deep Sleep State 28 Pentium® II Software Enhancements The Pentium® II and Pentium® III MSRs Instruction Set Changes Introduction Fast System Call/Return Instruction Pair Background The OS Initialization of the Fast Call Facility The OS Creates Four GDT Entries The OS Sets Up the Three MSRs The SYSENTER Instruction The SYSEXIT Instruction FP/SSE Save/Restore Instruction Pair Background Preparing for the Pentium® III’s Introduction of SSE When Executed on the Pentium® II Processor Detecting the FP/SSE Save/Restore Capability The FXSAVE Instruction The FXRSTOR Instruction The MXCSR Mask Field New/Altered Exceptions 29 Pentium® II Xeon Features Introduction To Avoid Confusion... Basic Characteristics Hardware Characteristics The Cartridge FSB Protocol Alteration (GTL+ to AGTL+) FSB Arbitration SMBus (System Management Bus) Note General SMBus Signals PSE-36 Mode PSE-36 Mode Background Detecting PSE-36 Mode Capability Enabling PSE-36 Mode Per Application Linear Memory Space = 4GB 386-Compatible Directory Lookup Mechanism Selected PDE Can Point to 4KB Page Table or a 4MB Page Linear Address Maps to a 4MB Page in 64GB Space Windows and PSE36 Part 9 Pentium® III 30 Pentium® III Hardware Overview One Product = Three Product Lines Pentium® II/Pentium® III Differences The Pentium® III/Xeon/Celeron Roadmap IOQ Depth The L1 Caches L1 Code Cache Characteristics L1 Data Cache Characteristics The L2 Cache The L2 Cache on the Early Pentium® III The Advanced Transfer Cache The Data Prefetcher SSE Introduced General Detecting SSE Capability Detailed Description of SSE The SSE Execution Units Introduction General The FP Multiplier Unit The Packed FP Add Unit The Shuffle/Logical Unit The Reciprocal/Reciprocal Square Root Unit Optimized Data Copy Operations 64-bit Paths Limited Performance The WCBs Were Enhanced Additional Writeback Buffers Background The Pentium® Pro and Pentium® II The Pentium® III SpeedStep Technology 31 Pentium® III Software Enhancements The Streaming SIMD Extensions (SSE) Why? Detecting SSE Support The SSE Elements The SSE Data Types General The 32-bit SP FP Numeric Format Background A Quick IEEE FP Primer The 32-bit SP FP Format Representing Special Values An Example Another Example Accuracy vs. Fast Real-Time 3D Processing The SSE Register Set The XMM Data Registers The MXCSR Loading and Storing the MXCSR Saving and Restoring the Register Set OS Support for FXSAVE/FXRSTOR, SSE and the SIMD FP Exception General Enable SSE/SSE2 and SSE Register Set Save and Restore Enable the SSE SIMD FP Exception SIMD (Packed) Operations Scalar Operations Cache-Related Instructions Overlapping Data Prefetch with Program Execution Streaming Store Instructions Introduction Some Questions Regarding Documentation The MOVNTPS Instruction The MOVNTQ Instruction The MASKMOVQ Instruction Ensuring Delivery of Writes Before Proceeding An Example Scenario The SFENCE Instruction Elimination of Mispredicted Branches Background SSE Misprediction Enhancements Comparisons and Bit Masks Min/Max Determination The Masked Move Operation Reciprocal and Reciprocal Square Root Operations MPEG-2 Motion Compensation Optimizing 3D Rasterization Performance Optimizing Motion-Estimation Performance Summary of the SSE Instruction Set SSE Alignment Checking The SIMD FP Exception SSE Setup CPUID Enhanced Serial Number Request Added Brand Index Request Added 32 Pentium® III Xeon Features Basic Characteristics PAT Feature (Page Attribute Table) What’s the Problem? Detecting PAT Support PAT Allows More Memory Types Default Setting of the IA32_CR_PAT MSR Entries Memory Type When Page Definition and MTTR Disagree General The UC- Memory Type Changing the Contents of the IA32_CR_PAT MSR Ensuring IA32_CR_PAT and MTRR Consistency Assigning Multiple Memory Types to a Single Physical Page Compatibility with Earlier IA32 Processors Part 10 Pentium® 4 33 Pentium® 4 Road Map The Roadmap 34 Pentium® 4 System Overview General The Graphics Adapter Device Adapters Snooping General A Memory Access Initiated by a Processor A Memory Access Initiated by a Device Adapter Definition of a Cluster Definition of the Boot Strap Processor The P6 Family BSP Selection Process The Pentium® 4 Family BSP Selection Process Starting up the Application Processors (the APs) 35 Pentium® 4 Processor Overview The Pentium® 4 Processor Family Pentium® III/Pentium® 4 Differences Pentium® 4/Pentium® 4 Prescott Differences Pentium® 4 Processor Basic Organization The FSB is Tuned for Multiprocessing Intro to the FSB Enhancements IA Instructions Vary in Length and Are Complex The Trace Cache There Are Two Pipeline Sections The µop Pipeline Introduction The P6 Processor’s Instruction Pipeline The Pentium® 4’s µop Pipeline The 90nm Pentium® 4’s Instruction Pipeline The IA32 Data Register Set Was Small General The P6 Had 40 General-Purpose Registers The Pentium® 4 Implements a Large Array of Data Registers The Compiler Manages Data Register Usage Elimination of False Register Dependencies Speculative Execution 36 Pentium® 4 PowerOn Configuration Configuration on Trailing-Edge of Reset Setup and Hold Time Requirements Built-In Self-Test (BIST) Trigger Assignment of IDs to the Processor Introduction The Cluster ID The Purpose of the Cluster ID The Cluster ID Assignment The Agent ID The Purpose of the Agent ID Physical versus Logical Processor The Agent ID Assignment Example Xeon MP System with Hyper-Threading Disabled Example Xeon MP System with Hyper-Threading Enabled Dual Processor System with Hyper-Threading Enabled A Single-Processor System with Hyper-Threading Enabled The Local APIC ID The Purpose of the Local APIC ID The Local APIC ID Assignment Error Observation Options In-Order Queue Depth Selection Power-On Restart Address Tri-State Mode Processor Core Speed Selection Bus Parking Option Description Bus Parking Configuration Hyper-Threading Option Program-Accessible Startup Features 37 Pentium® 4 Processor Startup Introduction The Processor’s State After Reset EAX, EDX Content After Reset Removal The Core Is Starving and Caching is Disabled Boot Strap Processor (BSP) Selection Introduction The BSP Selection Process How the APs are Discovered and Configured AP Detection and Configuration Introduction The BIOS’s AP Discovery Procedure Uni-Processor OS MP OS The FindAndInitAllCPUs Routine 38 Pentium® 4 Core Description One µop Doesn’t Necessarily = One IA32 Instruction Upstream vs. Downstream Introduction The Big Picture The Front-End Pipeline Stages CS:EIP Address Generation Linear to Physical Address Translation The L2 Cache Lookup On an L2 Miss, the Request Is Passed to the BSQ The Code Block Is Placed in the Instruction Streaming Buffer The Front-End BTB The Static Branch Predictor The Travels of a Conditional Branch Instruction The IA32 Instruction Decoder The P6 Instruction Decoder Was Complex The Pentium® 4 Decoder Is Simple The Trace Cache Can Keep Up with the Fast Execution Engine µops Are Streamed into the Trace Cache and the µop Queue Complex Instructions Are Decoded by the Microcode Store ROM General The MS ROM and Interrupts or Exceptions The Trace Cache General Build Mode Deliver Mode Self-Modifying Code (SMC) Introduction Your Code May Appear to be SMC SMC and the Earlier IA32 Processors SMC and The Pentium® 4 The Trace Cache BTB and the Return Stack Buffer The Trace Cache BTB The Return Stack Buffer (RSB) The µop Queue Intro to the µop Pipeline General The TC Next IP Stage The TC Fetch Stage The Drive 1 Stage The Allocator Stage The Register Rename Stage The Memory and General µop Queue Stage The Scheduler Stage The µop Dispatch Stage The Register File Stage The Execution Stage The Flags Stage The Branch Check Stage The Drive 2 Stage The µop Pipeline’s Major Elements The Allocator General The ReOrder Buffer (ROB) Entry The Register File Allocation The Load and Store Buffer Allocation The Memory or General µop Queue Allocation The Register Rename Unit General Renaming the Destination Register Renaming the Source Register(s) Renaming Eliminates False Register Dependencies The Memory and General µop Queues The Schedulers Enable Out-of-Order Execution The Register Files Are Strategically Placed Dispatch Port 0 Dispatch Port 1 Dispatch Port 2 Dispatch Port 3 Instruction Dispatch Rate The Complex Execution Units Are Pipelined The Retirement Stage General µop Retirement vs. IA32 Instruction Retirement Additional, Core-Specific Terms 39 Hyper-Threading General Background Multithreading Overview How Threads Are Assigned in an SMP System CMP Is Another Solution Traditional Single-Processor Multithreading The HT Approach Instruction Level Parallelism (ILP) But What If... This Requires Two, Almost Complete Register Sets HT = Simultaneous Multithreading Terms: Cluster, Physical CPU, Logical CPU Detecting HT Capability Enabling/Disabling HT Each Logical Processor Has Its Own Local APIC HT Processor Resource Types General Resources that Are Always Replicated Resources that Are Always Shared Resources Wherein Sharing or Replication Is Design-Specific The HT States Switching HT States Processor Enumeration The Primary and Secondary Logical Processor OS Support for HT General OSs that Include Native HT Support OSs that Are Compatible with HT OSs with No HT Support Overview of HT Resource Usage TC Access L2 Cache Access Code Block Is Placed in the Prefetch Streaming Buffer Instruction Decode Complex Instruction Decode The µops Are Placed in the Trace Cache The Return Stack Buffer Is Replicated The µops Are Placed in the µop Queue In Each Clock, the Allocator Switches Queues The Register Rename Stage µop Queues Are Partitioned The Schedulers Are Agnostic Register File Access The Retirement Stage HT and the Data TLB HT and the FSB The IOQ Depth Was Increased HT Performance Issues Introduction Thread Distribution to Logical Processors Load Balancing HT and the Processor Caches Physical Processors Operating on Separate Data Sets Data Sharing by Physical Processors Introduction Using a Semaphore to Access a Shared Data Area An Ideal Situation A Bad Situation If the Shared Data and the Semaphore Are in the Same Line Solution Data Sharing by Co-Resident Logical Processors Co-Resident Logical Processors with Separate Data Sets Executing Identical Threads Halt Usage Thread Synchronization Definition The Problem The Fix When A Thread Is Idle Spin-Lock Optimization WCB Usage HT and Serializing Instructions HT and the Microcode Update Feature HT Cache-Related Issues HT and the TLBs HT and the Thermal Monitor Feature HT and External Pin Usage STPCLK# Pin LINT0 AND LINT1 Pins A20M# Pin 40 The Pentium® 4 Caches A Cache Primer The L0 Cache Upstream vs. Downstream Overview Determining the Processor’s Cache Sizes and Structures Enabling/Disabling the Caches The L1 Data Cache General The L1 Data Cache Clients The Data Cache Is a Write-Through Cache The Data Cache is Non-Blocking Earlier Processor Caches Blocked, but So What The L1 Data Cache is Non-Blocking, and That’s Important! The L1 Data Cache Implements Squashing The L1 Data Cache Architecture The Data Cache’s View of Memory Space The Data Cache Lookup The Line Number Selects the Directory Set Simultaneously, a DTLB Lookup Is Performed The Physical Page Address Formation The Physical Page Address Compare The Data Cache LRU Algorithm The Data TLB (DTLB) The L2 ATC Introduction The L2 Cache’s Clients The L2 Cache Architecture The L2 Cache Is Non-Blocking The L2 Cache Implements Squashing The L2 Cache’s View of Memory Space The L2 Cache Lookup The L2 Cache LRU Algorithm General When an L2 Directory Entry Already Exists When an L2 Directory Entry Doesn’t Already Exist If There Is an L3 Cache If There Isn’t an L3 Cache Recording the New Sector in the L2 Cache Loads and TC Requests and the L2 Cache Stores and the L2 Cache Snoops and the L2 Cache Other L2 Cache Sizes The Hardware Data Prefetcher Introduction The Startup Penalty How the Data Prefetch Logic Works Some Constraints The L3 Cache Introduction The L3 Cache’s Client The L3 Cache Architecture The L3 Cache Is Non-Blocking The L3 Cache Implements Squashing The L3 Cache’s View of Memory Space The L3 Cache Lookup The L3 Cache LRU Algorithm General When an L3 Directory Entry Already Exists When an L3 Directory Entry Doesn’t Already Exist Loads and TC Requests and the L3 Cache Stores and the L3 Cache Snoops and the L3 Cache Other L3 Cache Sizes FSB Transactions and the Caches Background A Single-Sector Fetch A Two Sector Fetch Writeback of a Modified Line The Cache Management Instructions 41 Pentium® 4 Handling of Loads and Stores The Memory Type Defines Load/Store Characteristics Load µops The Load Buffers Loads from Cacheable Memory Loads Can Be Executed Out-of-Order The L1 Data Cache Implements Squashing Loads from Uncacheable Memory The Definition of a Speculatively Executed Load Replay Replay of µops Dependent on a Load Replay of Loads Dependent on a Store Loads and the Prefetch Instructions The LFENCE Instruction General LFENCE Ordering Rules Store-to-Load Forwarding Background Description Linear Address Mismatch Allows Load Before Store Linear Address Match Results in Store Forwarding Store Forwarding Rules Store µops Stores Are Handled by the Store Buffers Stores to UC Memory General UC Store Buffer Draining UC FSB Transactions Stores to WC Memory Determining if the WC Memory Type Is Supported The WC Memory Model WCB Evolution Filling the WCBs Draining the WCBs General Serializing Instructions A Special Use of the WCBs The WCBs and Hyper-Threading WCB FSB Transactions Stores to WP Memory General WP Store Buffer Draining WP FSB Transactions Stores to WT Memory General WT Store Buffer Draining Forcing a Buffer Drain The SFENCE Instruction General SFENCE Ordering Rules Sharing Access to a UC, WC, WP or WT Memory Region Stores to WB Memory Out-of-Order String Stores Stores and Hyper-Threading The MFENCE Instruction Non-Temporal Stores 42 The Pentium® 4 Prescott Introduction Increased Pipeline Depth Trace Cache Improvements Increased Trace Cache BTB Size Enhanced Trace Cache µop Encoding Increased Number of WCBs L1 Data Cache Changes Increased L2 Cache Size Enhanced Branch Prediction Enhanced Static Branch Predictor Dynamic Branch Prediction Enhanced Store Forwarding Improved Increased Number of Store Buffers Improved Load/Store Scheduling Force Forwarding Background Force Forwarding The Solution to False Forwarding The Address Misalignment Solution SSE3 Instruction Set Introduction Improved x87 FP-to-Integer Conversion Instruction The Problem The Solution New Complex Arithmetic Instructions Improved Motion Estimation Performance The Problem The Solution The Downside Instructions to Improve Processing of a Vertex Database Thread Synchronization Instructions Increased Elimination of Dependencies Enhanced Shifter/Rotator Integer Multiply Enhanced Scheduler Enhancements Fixed the MXCSR Serialization Problem Data Prefetch Instruction Execution Enhanced Improved the Hardware Data Prefetcher Hyper-Threading Improved Decreased Possibility of L1 Data Cache Blocking Increased the Size of the µop Queue Eliminated Page Table Walk/Split Line Access Conflict Handling Multiple Page Table Walks that Miss All Caches Trace Cache Responds Quicker to a Thread Stall The Data Cache and Hyper-Threading Author’s Note Introduction Shared Mode Adaptive Mode The MONITOR and MWAIT Instructions Background The Monitor Instruction The Mwait Instruction Example Code Usage The Wake Up Call 43 Pentium® 4 FSB Electrical Characteristics Introduction The Bus and Processor Clocks The BSEL Outputs The Processor’s Operational Clock Frequency BCLK Is a Differential Signal The Address and Data Strobes Delivering the Request The P6 Request Delivery Method The Pentium® 4/M Request Delivery Method Delivering the Data The P6 Data Delivery Method The Pentium® 4/M Data Delivery Method Why Multiple Strobes? The Data Bus Inversion Signals Address and Data Strobe Setup and Hold Specs The Voltage ID Everything’s Relative All AGTL+ Signals Are Active When Low All AGTL+ Signals Are Terminated Deasserting an AGTL+ Signal Line Each AGTL+ Input Has a Comparator The Reference Voltage The Sample Point The Pre-90nm Comparison The 90nm Comparison AGTL+ Setup and Hold Specs Signals that Can Be Driven by Multiple FSB Agents Minimum One BCLK Response Time 44 Intro to the Pentium® 4 FSB Enhanced Mode Scaleable Bus FSB Agents Agent Types Multiple Personalities Uniprocessor vs. Multiprocessor Bus The Request Agent The Request Agent Types The Agent ID The Purpose of the Agent ID How the Agent ID Is Assigned The Transaction Phases The P6 Transaction Phases The Pentium® 4/M Transaction Phases Transaction Pipelining The FSB Is Subdivided into Signal Groups Step 1: Gain Ownership of the Request Phase Signal Group Step 2: Issue the Transaction Request Step 3: Yield Request Phase Signal Group, Proceed to Next Signal Group The Phases Proceed in a Predefined Order The Request Phase The Snoop Phase The Response Phase The Data Phase(s) The Next Agent Can’t Use a Signal Group Until the Current Agent Is Finished With It Transaction Tracking Request Agent Transaction Tracking Snoop Agent Transaction Tracking Response Agent Transaction Tracking The IOQ 45 Pentium® 4 CPU Arbitration The Request Phase Logical versus Physical Processors The Discussion Assumes a Quad Xeon MP System Symmetric Agent Arbitration-Democracy at Work No External Arbiter Required The Arbitration Algorithm One Arbiter Per Physical Processor The Rotating ID The Busy/Idle Indicator General Reset’s Effect on the Busy/Idle Indicator The Idle Loop Transition from Idle to Busy Bus Parking Preemption by Another Physical Processor Transitioning Back to the Idle State Requesting Ownership Introduction Example of One Symmetric Agent Requesting Ownership Example of Two Symmetric Agents Requesting Ownership Definition of an Arbitration Event Once BREQn# Asserted, Keep Asserted Until Ownership Attained Example Case Where Transaction Cancelled Before Started 46 Pentium® 4 Priority Agent Arbitration Priority Agent Arbitration Example Priority Agents Priority Agent Beats Symmetric Agents, Unless... Using Simple Approach, Priority Agent Suffers Penalty Smarter Priority Agent Gets Ownership Faster Ownership Attained in 1 BCLK Ownership Attained in 2 BCLKs Be Fair to the Common People Priority Agent Parking 47 Pentium® 4 Locked Transaction Series Introduction The Shared Resource Concept Testing the Availability of and Gaining Ownership of Shared Resources A Race Condition Can Present a Problem Guaranteeing the Atomicity of a Read/Modify/Write The LOCK Instruction Prefix The Processor Automatically Asserts LOCK# on Some Operations Use Locked RMW to Test and Set a Semaphore The Duration of a Locked Transaction Series Back-to-Back RMW Operations Locking a Cache Line The Advantage of Cache Line Locking A New Directory Bit-Cache Line Locked The Memory Read and Invalidate Transaction (RWITM, or RFO) Line Containing a Semaphore Is in the E or M State Line Containing a Semaphore Isn’t in the L1 or L2 Cache Line Containing a Semaphore Is in the L2 Cache in the E State Line Containing a Semaphore Is in the Cache in the S State Line Containing a Semaphore Is in the Cache in the M State Semaphore Straddles Two Cache Lines 48 Pentium® 4 FSB Blocking Blocking New Requests-Stop! I’m Full! Assert BNR# When One Entry Remains BNR# Can Be Used by a Debug Tool Who Monitors BNR#? BNR# is a Shared Signal The Stalled/Throttled/Free Indicator Initial Entry to the Stalled State The Throttled State The Free State As an Agent Approaches Full, It Signals BNR# to Stall Everyone BNR# Behavior at Powerup BNR# Behavior During Runtime 49 Pentium® 4 FSB Request Phase Cautionary Note Introduction to the Request Phase The Source Synchronous Strobes The Request Phase Parity Request Phase Parity Checking ChipSet Request Phase Parity Checking and Reporting Processor Request Phase Parity Checking and Reporting The Request Phase Signal Group is Multiplexed Introduction to the Transaction Types The Contents of Request Packet A Description 32-bit vs. 36-bit Addresses The Contents of Request Packet B 50 Pentium® 4 FSB Snoop Phase Agents Involved in the Snoop Phase The Snoop Phase Has Two Purposes The Snoop Result Signals are Shared, DEFER# Isn’t The Snoop Phase Duration Is Variable There Is No Snoop Stall Duration Limit Memory Transaction Snooping The Snoop’s Effects on Processor Caches Self-Snooping Non-Memory Transactions Have a Snoop Phase 51 Pentium® 4 FSB Response and Data Phases A Note on Deferred Transactions The Purpose of the Response Phase The Response Phase Signal Group The Response Phase Start Point The Response Phase End Point The Response Types The Response Phase May Complete a Transaction The Data Phase Signal Group Five Example Scenarios A Transaction that Doesn’t Transfer Data A Read that Doesn’t Hit a Modified Line and is Not Deferred The Basics A Detailed Description How Does the Response Agent Know the Transfer Length? The Earliest Deassertion of DBSY# Special Case-Single BCLK, 0-Wait State Transfer A Write that Doesn’t Hit a Modified Line and Isn’t Deferred Introduction Transaction 1’s Response Transaction 1’s Target Is Ready to Accept Write Data Transaction 1’s Request Agent Gets the Go-Ahead Condition that Permits 1 BCLK TRDY# Assertion Transaction 1’s Request Agent Takes Ownership of the Data Bus Transaction 1’s Response Agent Drives Its Response Transaction 1’s Completion Transaction 2’s Description A Hard Failure Response The Snoop Agents Change the State of the Line from E to I or S to I A Read that Hits a Modified Line The Basics Relaxed DBSY# Deassertion A Write that Hits a Modified Line Data Phase Wait States The Response Phase Parity General ChipSet Response Phase Parity Checking and Reporting Processor Response Phase Parity Checking and Reporting Data Bus Parity Introduction ChipSet Data Phase Parity Checking and Reporting Processor Data Phase Parity Checking and Reporting Parity When Transferring a Sub-Block 52 Pentium® 4 FSB Transaction Deferral Example System Models Example Multi-Cluster Model The Problem Example Problem 1 Example Problem 2 Possible Solutions Example Read From a PCI Express Device The Read Receives the Deferred Response The Root Complex Performs the Read The Root Complex Issues a Deferred Reply Transaction General The Original Request Agent Is Selected The Root Complex Provides the Snoop Result Role Reversal in the Response Phase The Deferred Reply’s Data Phase All Trackers Retire the Transaction Other Possible Responses Example Write To a PCI Express Device The Write Receives the Defer Response The Root Complex Delivers the Write Data to the Target The Root Complex Issues a Deferred Reply Transaction General The Original Request Agent Is Selected The Root Complex Provides the Snoop Result Role Reversal in the Response Phase There is No Data Phase All Trackers Retire the Transaction Pentium® 4 Support for Transaction Deferral 53 Pentium® 4 FSB IO Transactions Introduction The IO Address Range The Data Transfer Length Behavior Permitted by the Spec How the Pentium® 4 Processor Operates 54 Pentium® 4 FSB Central Agent Transactions Point-to-Point vs. Broadcast The Interrupt Acknowledge Transaction Background The Transaction Details The Root Complex is the Response Agent The Special Transaction General The Message Types The BTM Transaction Is Used for Program Debug The Problem The Solution Enabling BTM Capability The BTM Transaction Packet A Composition Packet B Composition The Proper Response The Data Composition 55 Pentium® 4 FSB Miscellaneous Signals The Signals 56 Pentium® 4 Software Enhancements The Foundation Miscellaneous New Instructions General The Cache Line Flush Instruction The Fence Instructions The Memory Fence Instruction The Load Fence Instruction The Non-Temporal Store Instructions Introduction The MOVNTDQ Instruction The MOVNTPD Instruction The MOVNTI Instruction The MASKMOVDQU Instruction General When a Mask of All Zeros Is Used The PAUSE Instruction The Branch Hints Enhanced CPUID Instruction The SSE2 Instruction Set General DP FP Number Representation Packed and Scalar DP FP Instructions SSE2 64-Bit and 128-Bit SIMD Integer Instructions SSE2 128-Bit SIMD Integer Instruction Extensions Your Choice: Accuracy or Speed The SSE3 Instruction Set Local APIC Enhancements The Thermal Monitoring Facilities Introduction to Thermal Monitoring Thermal Monitor Feature Detection Stop Clock Acts as a Gate for the Processor Clock Catastrophic Shutdown Detector Automatic Thermal Monitoring Thermal Monitor and Interrupts Interrupts Are Blocked While Stop Clock Is Low The Thermal Monitor Interrupt Software Controlled Clock Modulation Relationship of the Hardware- and Software-Based Mechanisms HTT and Thermal Monitoring FPU Enhancement General Fopcode Compatibility Mode The MSRs The Machine Check Architecture Introduction The Pentium® 4 MCA Enhancements The Extended MC State MSRs Last Branch, Interrupt, and Exception Recording The Debug Store (DS) Mechanism Introduction Feature Detection Setting Up the DS Feature Enabling the BTS Feature Enabling the PEBS Feature The PEBS Record Format The BTS Record Format New Exceptions The Performance Monitoring Facility Performance Monitoring Is Not Architecturally Defined Author’s Note An Overview There Are Two Event Categories There Are Three Sampling Methods Relationship of a Counter, Its CCCR and the ESCRs The Event Select Control Registers The Counter Configuration Control Registers General Counter Cascading Interrupt on Overflow Extended Cascading Accessing the Performance Counters Halting Event Counting Non-Retirement Event Counting Introduction The Set Up The Event Filtering Mechanism Introduction Threshold Comparison The Threshold Condition Transition Filter At-Retirement Event Counting First, Some Terminology Bogus, Non-Bogus, Retire Tagging Replay Assist General The Tagging Mechanisms Introduction Multi-Tagging PEBS and Multi-Tagging Some µops Cannot Be Tagged Front-End Tagging Execution Tagging Replay Tagging Precise Event-Based Sampling General Limited To a Single Counter Detecting the PEBS Capability Enabling PEBS The PEBS Interrupt Handler Sometimes, the DS Feature Is Disabled PEBS and Hyper-Threading Counting Clocks Introduction The Non-Halted Clockticks Measurement The Non-Sleep Clockticks Measurement The Time Stamp Counter 57 Pentium® 4 Xeon Features General The Pentium® 4 Xeon DP The Pentium® 4 Xeon MP Part 11 Pentium® M 58 Pentium® M Processor Background The Pentium® M and Centrino Characteristics Overview The FSB Characteristics Uses the Pentium® 4 FSB Protocol Pentium® M-Specific Signals FSB Power Utilization Enhancements Enhanced Power Management Characteristics Background Entry to the Deep Sleep State The Deeper Sleep State Enhanced SpeedStep Background Enhanced SpeedStep Description Three Different Packaging Models Improved Thermal Monitor Mode Enhanced Branch Prediction Introduction The Loop Detector The Indirect Branch Predictor The Problem Indirect Branch Predictor Description µop Fusion Background µop Fusion Description General The Fused Store The Fused Load and Operate Advanced Stack Management Background Advanced Stack Management Description Miscellaneous Hardware-Based Data Prefetcher The L2 Cache The Data Cache and Hyper-Threading The Next Pentium® M Part 12 Additional Topics 59 CPU Identification Prior to the Advent of the CPUID Instruction Determining if the CPUID instruction Is Supported General Determining the Request Types Supported Determining Basic Request Types Supported Determining Extended Request Types Supported The Basic Request Types Request Type 1 General The Brand Index Request Type 2 General Request for Cache and TLB Information Request Type 3 Request Type 4 Request Type 5 The Extended Request Types Enhanced Processor Signature 60 System Management Mode (SMM) What Falls Under the Heading of System Management? The Genesis of SMM SMM Has Its Own Private Memory Space The Basic Elements of SMM A Very Simple Example Scenario How the Processor Knows the SM Memory Start Address Protected Mode, Paging and PAE-36 Mode Are Disabled The Organization of SM RAM Entering SMM The SMI Interrupt Is Generated No Interruptions Please General Exceptions and Software Interrupts Permitted but Not Recommended Servicing Maskable Interrupts While in the Handler Single-Stepping through the SM Handler If Interrupts/Exceptions Permitted, Build an IDT SMM Uses Real Mode Address Formation NMI Handling While in SMM Default NMI Handing How to Re-Enable NMI Recognition in the SM Handler If an SMI Occurs within the NMI Handler Informing the Chipset That SMM Has Been Entered General A Note Concerning Memory-Mapped IO Ports The Context Save General Although Saved, Some Register Images Are Forbidden Territory Special Actions Required on a Request for Power Down The Register Settings on Initiation of the SM Handler The SMM Revision ID The Body of the Handler Exiting SMM The Resume Instruction Informing the Chipset That SMM Has Been Exited The Auto Halt Restart Feature Executing the HLT Instruction in the SM Handler The IO Instruction Restart Feature Introduction An Example Scenario The Detail Back-to-Back SMIs During IO Instruction Restart Caching from SM Memory Background The Physical Mapping of SM RAM Accesses FLUSH# and SMI# Description A Cautionary Note Regarding the Pentium® Setting Up the SMI Handler in SM Memory Relocating the SM RAM Base Address Description In an MP System, Each Processor Must Have a Separate State Save Area Accessing SM RAM Above the First MB SMM in an MP System 61 The Local and IO APICs Before the Advent of the APIC MP Systems Need a Better Interrupt Distribution Mechanism Introduction The APIC Interrupt Distribution Mechanism Introduction Message Transfer Mechanism Prior to the Pentium® 4 Message Transfer Mechanism Starting with the Pentium® 4 Inter-Processor Interrupt Messages Local Interrupts NMI, SMI and Init Messages The Cluster and APIC ID Physical Destination Mode Logical Destination Mode A Short History of the APIC The Introduction of the APIC Pentium® Pro APIC Enhancements The Pentium® II and Pentium® III Pentium® 4 APIC Enhancements Detecting the Presence and Version of the Local APIC Enabling/Disabling the Local APIC General Permanently Disabling the Local APIC Temporarily Disabling the Local APIC Operational Characteristics of a Disabled Local APIC Local Cluster and APIC ID Assignment Cluster ID Assignment APIC ID Assignment Maximum Number of Local APICs BIOS/OS Reassignment of Local APIC IDs The Local APIC IDs Are Stored in the MP and ACPI Tables Reading the Local APIC ID An Introduction to the Interrupt Sources Local Interrupt Sources Remote Interrupt Sources Introduction to Interrupt Priority General Definition of a User-Defined Interrupt The Priority Amongst the User-Defined Interrupts Definition of Fixed Interrupts Masking User-Defined Interrupts An Intro to Edge-Triggered Interrupts An Intro to Level-Sensitive Interrupts The Local APIC Register Set Local and IO APIC Register Areas Are Uncacheable Introduction to the Local APIC’s Register Set The IRR, TMR and ISR Registers General An Example The EOI and Its Effects Interrupt Request Buffering Register Access Alignment Locally Generated Interrupts Introduction The Local Vector Table The Pentium® Family’s LVT The P6 Family’s LVT The Pentium® 4 Family’s LVT Local Interrupt 0 (LINT0) Introduction The Mask Bit The Trigger Mode and the Input Pin Polarity The Delivery Mode The Vector Field The Remote IRR Bit The Delivery Status Local Interrupt 1 (LINT1) The Local APIC Timer General The Divide Configuration Register One Shot Mode Periodic Mode The Performance Counter Overflow Interrupt The Thermal Sensor Interrupt The Local APIC’s Error Interrupt Task and Processor Priority Introduction The Task Priority Register (TPR) The Processor Priority Register (PPR) The User-Defined Interrupt Eligibility Test Interrupt Messages Introduction Sending a Message From the Local APIC Physical Destination Mode Logical Destination Mode Introduction The Flat Model The Cluster Model The Flat Cluster Model The Hierarchical Cluster Model Lowest-Priority Delivery Mode General Chipset-Assisted Lowest-Priority Delivery The IO APIC The Purpose of the IO APIC Overview of an Edge-Triggered Interrupt Delivery Overview of a Level-Sensitive Interrupt Delivery The IO APIC Register Set The IO APIC Register Set Base Address The IO APIC Register Set The IRQ Pin Assertion Register The EOI Register Non-Shareable IRQ Lines Shareable IRQ Lines Linked List of Interrupt Handlers How It Works The ID Register The Version Register The Redirection Table Register Set Interrupt Delivery Order Is Rotational Message Signaled Interrupts (MSI) General Using the IO APIC as a Surrogate Message Sender Direct-Delivery of the MSI Memory Already Sync’d When Interrupt Handler Entered The Problem The Old Way of Solving the Problem How MSI Solves the Problem Message Format The FSB Message Format The APIC Bus Message Format The Spurious Interrupt Vector The Problem The Solution Additional Spurious Vector Register Features The Agents in an Interrupt Message Transaction MCH Initiates an Interrupt Message Transaction A Local APIC Initiates an Interrupt Message Transaction BSP Selection Process Introduction to the BSP Selection Process The P6 Family BSP Selection Process The Pentium® 4 Family BSP Selection Process The APIC, the MPS and ACPI Acronyms Index Symbols Numerics A B C D E F G H I J K L M N O P Q R S T U V W X Z
Similar books
MySQL® Notes for Professionals book
2018 · PDF
MrExcel 2022: Boosting Excel
2022 · PDF
MrExcel 2022: Boosting Excel
2022 · PDF
Session C11: Ancient Cultural Landscapes in South Europe – their Ecological Setting and Evolution, Session C22: Gardeners from South America, Session S04: Agro-Pastoralism and Early Metallurgy Sessions, Session WS29: The Idea of Enclosure in Recent Iberian Prehistory, Session C88: Rhytmes et causalites des dynamiques de l'anthropisation en Europe entre 6500 ET 500 BC: Hypotheses socio-culturelles et/ou climatiques: Proceedings of the XV UISPP World Congress (Lisbon 4-9 September 2006) / Actes du XV Congrès Mondial (Lisbonne 4-9 Septembre 2006) Vol.36
2010 · PDF
THE BRITISH ARMY IN INDIA: ITS PRESERVATION BY AN APPROPRIATE CLOTHING, HOUSING, LOCATING, RECREATIVE EMPLOYMENT, AND HOPEFUL ENCOURAGEMENT OF THE TROOPS. with AN APPENDIX ON INDIA : THE CLIMATE OP ITS HILLS ; THE DEVELOPMENT OF ITS RESODRCBS, INDUSTRY, AND ARTS ; THE ADMINISTRATION OF JUSTICE ; THE BLACK ACT ; THE PROGRESS OF CHRISTIANITY ; THE TRAFFIC IN OPIUM ; THE VALUE OF INDIA ; PERMANENT CAUSES OF DISAFFECTION, AND OF THE RECENT REBELLION ; THE TRADITIONARY POLICY; MISGOVERNMENT BY NATIVE RULERS ; ANNEXATIONS OF THEIR TERRITORY, ETC.
1858 · PDF
Idries Shah 27 Books Collection : A Perfumed Scorpion, A Veiled Gazelle, Caravan of Dreams, Darkest England, Destination Mecca, Evenings with Idries Shah, Knowing How to Know, Learning How to Learn, Letters and Lectures of Idries Shah, Neglected aspects of Sufi study, Observations, Oriental Magic, Reflections, Seeker after Truth, Special Illumination, Special Problems in the study of Sufi ideas, Sufi thought and action, Tales of the Dervishes, The Dermis Probe, The Elephant in the Dark, The Englishman Handbook, Idries Shah Antology, The Magic Monastery, The natives are restless, wisdom of the Idiots PDF.
2022 · PDF
The travels of Capts. Lewis and Clarke from St. Louis, by way of the Missouri and Columbia rivers, to the Pacific ocean; performed in the years 1804, 1805 & 1806, by order of the government of the United States. Containing delineations of the manners, customs, religion, &c. of the Indians, comp. from various authentic sources, and original documents, and a summary of the Statistical view of the Indian nations, from the official communication of Meriwether Lewis. Illustrated with a map of the country, inhabited by the western tribes of Indians
1809 · PDF
Professional Linux kernel architecture ''Wrox programmer to programmer''--Cover. - ''What you are reading right now is the result of an evolution over more than seven years: After two years of writing, the first edition was published in German by Carl Hanser Verlag in 2003. It then described kernel 2.6.0. The test was used as a basis for the low-level design documentation for the EAL4+ security evaluation of Red Hat Enterprise Linux 5, requiring to update it to kernel 2.6.18 (if the EAL acronym does not mean anything to you, then Wikipedia is once more your friend). Hewlett-Packard sponsored the translation into English and has, thankfully, granted the rights to publish the result. Updates to kernel 2.6.24 were then performed specifically for this book''--P. ix
2008 · PDF