Monday, August 8, 2011

Dell's Powerhouse R910 Review

We've been using the R910 now for about 10 months. During those 10 months, I've learned this is one of those workhorse servers that can take anything thrown at it and come back for more. We have 3 of them running as HyperV host machines, and 2 running Oracle 10g and 11g Enterprise loaded with quad 8-core Intel Xeon CPUs, 128gb RAM, Perc H700 Raid Controller, with 4 146gb 15k SaS drives in RAID10 for the O/S and 12 SSD's from various manufacturers (OWC and the Crucial M4) in another RAID10 array with dual Broadcom 10gbE nics. These machines would fly outta the cabinet if they weren't secured with the sliding versa-rails. The raw speed of any operation is just crazy. IDrac6 makes remote management a breeze compared to the previous Dracs. The ease of service on these is nice, but I don't think we'll be needing that anytime soon. My ideal box would be like these, but with multiple 600gb 15k SaS drives for lower volatility data sets. I would add-in some Fusion-io or Virident cards for my data sets requiring extreme I/O, low latency, and high throughput. Imaging a DML statement running in parallel across 32 cores reading & writing to several FIO or Virident cards simultaneously. This server platform has plenty of PCI-E slots to fill up with those types of cards. Yes, in my mind, the R910 is the beast under your bed, in your closet, or in your head....

- Posted using BlogPress from my iPad

Location:Columbus, Ga

Thursday, May 5, 2011

VOIP in the Small-Medium Business

Recently, I was faced with our aging legacy digital phone system coming off-lease. Years ago, when it was implemented, we looked at VOIP, but decided that it was too immature at the time. I planned to revisit it again when this system was to be replaced. That time has come and I looked at the major players: Cisco, Avaya, NEC, and Fonality/Trixbox.

After many weeks of evaluation and comparison, I realized they they all provide basically the same functionality. There was no big "killer feature" the differentiated one above the others, so I looked at "openess"and pricing. It seems that the big players like to charge extra for using SIP trunks, presence, and other options of the sort. Some vendors wanted 3 physical appliances or servers. Presence requires a dedicated server? I think that is to push you into spending more and adding more of their hardware into your shop. That's the feeling I got. One of the vendors wanted to charge per SIP trunk added to the system. I had enough of that sort of ridiculousness and looked deeply into the Trixbox/Fonality solution.

Trixbox Pro is basically a customized CentOS Linux 4.4 installation with the Trixbox Pro modules. I simply paid the licensing fees, downloaded the ISO image, burned it to CD, inserted it into my server, and booted from it. It looked like a normal CentOS install and asked for a bit of info such as root password, hostname, IP address, timezone, etc. Easy to follow. Afterwards, the server rebooted and I logged in, activated the server with Fonality, and connected to the management console. To my surprise, the management console is hosted by Fonality and trickles changes down to your server. I wasn't too thrilled with that, but it works. The Community Edition does not work like that. It is all hosted on your end. All clients connect to your local server and remote clients go thru Fonality to connect down to your server, again unlike the Community Edition. It took me about a day to add all 50 of my users and conference bridges. It seems like almost any SIP-compliant handset or softphone will work with it. We are pleased with it so far.

This is what made the decision easy to me:

Cisco Solution: $75,000-$86,000
Avaya Solution: $45,000-$65,000
NEC Solution: $46,000-$64,000
Trixbox Pro Solution: <$18,000

Simply, I got basically the same functionality with MORE flexibility for FAR less money with lifetime updates included. Professional support is only $2,000 per year as well. There really wasn't a decision to make after looking at it from a high-level overview. 

Friday, April 1, 2011

It's in the "Clooooud"....

Haha, that phrase just cracks me up. Hearing the typical "Web 2.0" fanboi belch out that phrase just makes me want to go "Falling Down". They say "Cloud" like its something profoundly new and great. While it may be great in some cases, it sure as heck ain't new. Let's break it down.... does a "Cloud" not have:


  • Datacenter Floorspace?
  • Server Cabinets or Racks?
  • Servers?
  • CPU's?
  • RAM?
  • Disk Storage?
  • Network Interconnects?
  • UPS?
  • Generator backups?
  • Windows, Linux, or Unix Operating Systems?
  • Shared resources?

Really? Are you sure? Of course they do. Don't stutter with glazed eyes and just say "It's the cloooud man.." It is the same things we've been doing for years. Same pig, different sauce. Its another buzz-word thought up by some ITIL/Six-Sigma/Business Analyst-wannabe techie type that wanted to coin a new slang phrase to make their self feel important. Whoop de do.

For a multi-million dollar line of business application, why would you even consider using a cloud vendor? Cost? Simplicity? Of course not. Only because you let a developer-type make infrastructure decisions. That is an epic fail. Selecting a cloud vendor means you do NOT have physical access to your production data, servers, or network infrastructure. Business App offline? Call your cloud vendor and pray they can get the issue(s) resolved to meet your SLA requirements. If you don't have physical control over your infrastructure, you can't *REALLY* guarantee anything to your customers because you ARE at the mercy of your *Buzzword* Vendor.

Don't get me wrong, cloud computing is useful in certain scenarios, but for hosting mission-critical applications, it is insane in my mind. I'm sure I'll get a bunch of flaming rebuttals posted here from cloud vendors and cloud-developer fanbois, but remember, you drank the Kool-Aid, and this is MY blog :-)

Toodles,

-Me

Friday, March 11, 2011

Oracle 10g dbconsole configuration error

I noticed lately when I install Oracle 10g on Windows or Linux, when I try to install/configure Enterprise Manager Database Control (DBConsole), it errors out and will not complete. No matter what I tried, dropping & creating the repository over and over, still didn't work.

I stumbled across Oracle Metalink Patch 8350262, which supposedly fixes the issue. The issue is that the self-generated certificate expired on 12-31-2010. After installing 10g, apply Patch 8350262 (it is an opatch patch), everything went fine as usual. I hope this saves you some headache and time.

Saturday, February 12, 2011

SSD in the Enterprise

I recently wrote a high-level article focusing on SSD use in an Enterprise environment for my friend Les at TheSSDReview. It has some great info in it including use case scenarios. The article can be found here:

http://thessdreview.com/latest-buzz/ssd-migration-in-the-enterprise-environment/

The goal is to alleviate concerns with using SSD in an enterprise environment. Far too many companies could tremendously benefit from SSD use in many cases with the right planning, preparation, and expectations. I plan to write many more for him, so stay tuned.


Tuesday, January 25, 2011

Fusion-io vs. OCZ

I've seen a lot of interest in my past posts regarding the testing of the Fusion-io ioDrive and the OCZ Z-Drive. I felt that I needed to clarify my position and opinions on the two products a bit.

If you are an end-user or enthusiast and your livelihood doesn't depend on the data stored by the product, the OCZ Z-Drive would probably be the better product for you because of the cost difference between the two.

If you are a business user using the products in servers where critical data is stored, your ONLY option at that point is the Fusion-io products, no question. DO NOT trust your critical business data to the OCZ-ZDrive, do not even think about it. I experienced a failure firsthand within 3 days of receiving the one I ordered, and I returned it. Fusion-io has well engineered and backed products. They perform extremely well, and are not "built upon another product" like the OCZ product is.

Let me clarify that, the OCZ product is simply an LSI RAID controller with MLC or SLC NAND memory modules. The RAID controller BIOS only allows RAID0 and RAID1. With 8 banks of memory, that means with RAID0, ONE big drive with no redundancy in the event of failure. With RAID1, that means 4 logical drives in which to store your data. One NAND bank can fail and you're still in business. With RAID0, how it is out of the box, one failure = TOTAL data loss. You're out of business if that was critical data.

Fusion-io has consumer-level products for the enthusiast users. I'd tend to recommend it to those types whom doesn't need large high performance storage needs. If you need 200+gb, then the OCZ Z-Drive would make sense if the data stored isn't critical. If so, reconfigure the RAID BIOS to use RAID1 and divide your data across the 4 logical drives. You'll still get great performance, but not as much as it would be with 8-way RAID0 with the threat of data loss present.

If you are a business needing high performance data storage,  I urge you to look into Fusion-io and their products. I can think of no better performing and robust product on the market that the Fusion-io ioDrive, ioDrive Duo, or ioDrive Octal.

For shops whom need the absolute highest level of performance for their SQL Server, Oracle Database environments, or any other I/O intensive workloads, nothing can beat the Fusion-io Octal drive for the money. Yes, it is VERY expensive, but not to those companies whom need that level of performance. If it is too much for you, the ioDrive and ioDrive Duo is your best bang-for-the-buck choice. Hands-down, you will not be disappointed with the level of performance or support that you get from Fusion-io.

Friday, December 3, 2010

Tools of the trade

I just have to plug two of my favorite products here. They are absolutely great at what they do and make my life a LOT easier.

Confluence: This is an enterprise-grade Wiki for storing, categorizing, and searching documentation as well as collaboration. It is EASY to deploy, manage, and upgrade. Take a test drive of it.

Jira: This is a workflow tool allowing you to enter tasks & issues, track progress, and time spent on those. Our development team uses it for issue tracking and other things, but it is absolutely GREAT just like Confluence.

Wednesday, October 27, 2010

Google Enterprise Services

FYI: Do not buy ANY Google Enterprise Services. When you call for support, you get punted to their support forums, and the call handler will drop you after stating there is nothing they can do for you.

I was simply trying to ask how do I go about purchasing MORE seats for Postini, but with that attitude, and case-ownership attitude, I believe I will spend my company's money elsewhere.

Sunday, October 10, 2010

Arista Networks Host-Link Aggregation

Setup your server side (I use Broadcom Advanced Control Suite) as you normally would.

On your new shiny Arista 7120T, this is what you need to do:

7120t enable
config t
int et1-4
channel-group 1 mode on

Then run: show int po1
Port-Channel1 is up, line protocol is up
Hardware is Port-Channel, address is xxxx.xxxx.xxxx
MTU is 9212 bytes, BW 40000000 Kbit
Full-Duplex, 40Gb/s
Active Members in this channel: 4
...Ethernet1 , Full-duplex, 10Gb/s
...Ethernet2 , Full-duplex, 10Gb/s
...Ethernet3 , Full-duplex, 10Gb/s
...Ethernet4 , Full-duplex, 10Gb/s

5 minutes input rate 0 bits/sec (0%), 0 packets/sec
5 minutes output rate 0 bits/sec (0%), 0 packets/sec
   0 packets input, 0 bytes
   Received 0 broadcasts, 0 multicast
   0 input errors
   0 packets output, 0 bytes
   Sent 0 broadcasts




Now you're cooking with gas :)


For switch to switch LACP: 


7124-SWB(config)#int et1-4
7124-SWB(config-if-Et1-4)#channel-group 1 mode passive
7148-SWA(config)#int et1-4
7148-SWA(config-if-Et1-4)#channel-group 1 mode active

Once the above sequence is entered and synchronization is complete, the ethernet interfaces 1-4 on each switch are bound to the port-channel.

If you're using another switch with your Arista, the config might be a bit different and you may need to try a few different configs to get it to work correctly.

Friday, October 8, 2010

SSD Goodness

JACKPOT!

Interesting, OWC REALLY enters the fray now.

I can't wait to see Larry release THIS upon the world. My Oracle servers are already drooling over this as they watch me type this over my shoulders. Bring them on!

Monday, October 4, 2010

OWC SQLIO Testing

I just finished the SQLIO testing of the OWC Mercury Extreme Pro RE SSD's.

Platform Details: Dell PowerEdge 6950, 4x Dual-Core AMD Opteron 2.6ghz CPU's, 32GB RAM, Dell Perc 5/i RAID Adapter. RAID details: RAID10, 1MB Stripe Size, Adaptive Read Ahead, Writeback Cache, Windows Server 2008 R2 Standard. I used 4x of the 200GB version of the drive in RAID10 and installed WS2008R2 on that array.

The results were about what I was expecting. See below.






















































The full details are available here.

Thursday, September 30, 2010

OWC SSD Preliminary Testing

Just got in our large order of the OWC Mercury Extreme Raid Edition SSD's (200gb size). While I am waiting on the R910's they are destined for, I threw a few into an older Dell PowerEdge 6950 with quad Dual-Core AMD Opteron 2.6ghz CPU's. I setup a single boot drive and a 4-disk RAID10 group. The RAID settings were set to 1MB stripe size, adaptive read-ahead, and writeback. Went through a normal Windows 2008 R2 Standard installation to run my IOMeter and (soon) SQLIO tests.

The results were pretty good.

Using 4MB IO sizes, they REALLY shined bursting over 2800 MB/s of thruput.




















The details of the test were: 4MB IO size, 50% read, 50% write, 50% random and 50% sequential. The sustained average was a more-sane 1100 MB/s with the same settings.





















Using 4k IO's, I saw over 600 MB/s, which is about what I expected, and the IOPs were good at >50,000.

Not bad for ~$2400 of storage, with a full 5-year replacement warranty.

Stay tuned for further, more detailed testing including SQLIO. I know you SQL Server guys like to see those. I'm running Oracle 11g R2 Enterprise on these 4 right now as a sandbox testing environment. Using these SSD's with Oracle parallel execution has made a GIANT impact on our application performance so far. Keep in mind, we're not a production shop, and no important data was at risk. This is just a "testing the waters" scenario before we move further.

Wednesday, September 15, 2010

10GBase-T is here

You read that right, 10-Gigabit Ethernet over CAT5E or CAT6A, with distance limitations of course, and good cabling.

This is the Arista Networks 7120T-4S 10GBase-T switch. I am anxious to get this installed and tested with the new servers coming in multiple Broadcom 57710 NICs. Hyper-V and ESX should run nicely with 4x 10GBase-T in LACP per physical host.

Now for the pics:








































Sunday, September 5, 2010

OWC SSD in Macbook Pros

I received the two 13" MacBook Pro's and their OWC SSD's yesterday:




















Installation was a snap taking a couple of minutes each:




















After installing Mac OS X 10.6, I saw just how fast these things are. From the "dong" to logged in was only a matter of seconds, maybe 10 seconds from the power-on button.

I setup Bootcamp with Windows 7 X64 Ultimate because the two end-users requested this. After setting that up and installing the appropriate Bootcamp drivers, I updated them and rebooted.

I tried IOMeter  and saw good results, right at what was advertised. Being my curious self, I bumped the IO Size to 4MB and was surprised to see the performance even higher than advertised at 339MBs:





















Now I am looking forward to seeing what 12 of those in RAID10 and RAID0 will do. Can't beat that for the price.

Tuesday, August 31, 2010

Oracle Bigfile Tablespaces

After many years of ignoring Bigfile Tablespaces, I finally devoted some time to investigating them. I've been missing out. I've always hated building a "shell" database instance, then running my script to build my normal small file tablespaces, and add 10+ 8GB datafiles to each of my six main tablespaces. This is a LONG running script because the disk IO is not very fast in my existing SAN storage. This will not be an issue anymore in the near future, thanks to the folks at OWC. Still, managing 20-40 datafiles per tablespace is just plain aggravating. Now, I have one datafile per tablespace, that grows by 1GB in size as needed up to the filesystem limitation I am using, which is 32TB. I do not plan on outgrowing that anytime in the near future. Sure, it makes Rman recovery slower when recovering a single 300+gb datafile, but again, not on SSD. I don't use rman to begin with. Datapump thru bash shell scripts with Crontab is my backup of choice for my Linux/Oracle servers. Bigfile tablespaces will be going in during my upcoming server/storage refresh for my main tablespaces.

OWC Mercury Extreme Pro SSD's

The verdict is in. We're going with the OWC RE Pro SSD's in our Hyper-V and Oracle DB servers. Why, you ask? Well, the Raid Edition comes with a full 5-year warranty and 28% over-provisioning. My testing of 3x Corsair F240's convinced me the Sandforce-based SSD's are Enterprise-class, and they have the performance I want. Corsair doesn't have the 5-year warranty of the OWC product. OWC put it in writing that if I happen to "write-out" the drives before the warranty is up, they will replace them under warranty. That's standing by your product. I only plan to use them for 3 years because that's our leasing cycle. They'll be replaced in 3 years, so I'm not worrying about the write-cycle lifetime in RAID10. We're also putting them in all developer workstations and notebooks from here on out.

Wednesday, August 11, 2010

Direct I/O vs. Cached I/O

The simple difference between Direct I/O and Cached I/O:

Direct I/O :


































Cached I/O :



































Keep in mind, all of your data won't be cached all of the time. Only a small amount will be, so you won't get that high performance 100% of the time. If you want that, go with SSD.

Tuesday, August 10, 2010

Corsair F240 SSD

I received them yesterday. I installed one in my Mac Pro and copied my Win7 VM to it.

Just finished running ATTO and this was the result:





































To make formatting easier, here are the SQLIO results as a linked image:

MB/S
IOPS

Remember, this is a SINGLE drive in my Mac Pro, running a Windows 7 Ultimate X64 VMWare Fusion VM on it. The VM has 4GB of RAM and 4 CPU cores allocated to it. These tests were ran from within that VM. I know it wasn't "optimal" conditions, but the results look promising given the scenario. Not bad for a single $600 drive.



This is 3 of the F240's in RAID0:


































 Now that makes a bit more sense. Not bad at all.

Of course with NAND, TRIM doesn't work with RAID, so the life of the drive may be shortened somewhat until support for TRIM with RAID surfaces.

Friday, August 6, 2010

Oracle ASM

What a gigantic HEADACHE. I can see how it can be useful from an administrative standpoint and how it is tailored for RAC, but I am skeptical of the supposed "performance" increase. If it slightly increases performance, is it worth the headache it causes from a configuration and maintenance standpoint? To me, absolutely NOT. It takes a DAY or more to set it up and configure it correctly. This is BEFORE the DB software is installed/configured, and the database being built. Using regular SAN LUNS, I can partition, mount, and be ready to build the database in ~30 minutes. ASM is not worth the headache to me.

Why would I want a SECOND Oracle instance just to manage my DB storage disks? That's more CPU and memory being eaten up just for storage purposes.