Bienvenidos
Todos estos pasos descriptos fueron probados en ambientes productivos
martes, 30 de mayo de 2017
Para cambiar ip o setear la ip e un t5220 o t71/2
Primero configurar al SP ( Service Processor ) para que la Ip sea estatica
-> set /SP/network pendingdiscovery=static
Seteo la direccion IP para la SP
-> set /SP/network pendingipaddress=10.67.33.115
Setear la direccion IP del gateway
-> set /SP/network pendingipgateway=10.67.33.1
Setear el netmask para la SP
-> set /SP/network pendingipnetmask=255.255.255.0
*** las ip son a modo de ejemplo, tanto la de gateway como la del SP y netmask
Instalacion del Agente de Control M , Version 9.0.00 y su Fix Pack
Para la instalacion vamos a crear previamente el usuario que definimos con anterioridad, x300168 y el grupo controlm
El directorio de instalacion sera el mismo del usuario, para el caso que definimos es el /export/home/x300168
Los servidores de ControlM tienen como hostname controlm .
-rwxr-xr-x 1 root root 99M Jul 12 2016 PAKAI.9.0.00.200_Solaris_INSTALL.BIN
[sun5004] NUEVO /export/home/x300168 # ls -ltr
total 152
-rwxr-xr-x 1 root root 6656 Jun 9 2014 setup.sh
-rwxrwxrwx 1 root root 66064 Jun 8 2015 setup.exe
drwxr-xr-x 10 root root 512 Jun 9 2015 UNIX
drwxr-xr-x 4 root root 512 Jun 9 2015 Windows
drwxr-xr-x 5 root root 512 Jun 9 2015 FORMS
drwxr-xr-x 6 root root 512 Mar 17 10:33 OLD
[sun5004] NUEVO /export/home/x300168 # sh setup.sh
Starting graphical user interface...
The $DISPLAY environment variable is not defined.
Install will continue in non graphical mode.
Are you sure you want to continue?
==
Enter command:
=== Control-M/Agent 9.0.00 Installation ===
=== License Agreement ===
END USER LICENSE AGREEMENT
BY OPENING THE PACKAGE, INSTALLING, PRESSING "AGREE" OR "YES" OR USING THE PRODUCT, THE ENTITY OR INDIVIDUAL ENTERING INTO THIS AGREEMENT AGREES TO B
E BOUND BY THE FOLLOWING TERMS. IF YOU DO NOT AGREE WITH ANY OF THESE TERMS, DO NOT INSTALL OR USE THE PRODUCT, PROMPTLY RETURN THE PRODUCT TO BMC OR
YOUR BMC RESELLER, AND IF YOU RETURNED THE LICENSE WITHIN 15 DAYS OF THE DATE OF YOUR ORDER, CONTACT BMC OR YOUR BMC RESELLER FOR A REFUND OF LICENS
E FEES PAID. IF YOU REJECT THIS AGREEMENT, YOU WILL NOT ACQUIRE ANY LICENSE TO USE THE PRODUCT.
This Agreement ("Agreement") is between the entity or individual entering into this Agreement ("Customer") and the BMC Entity for the applicable Regi
on where Customer acquired the License as described in Section 18 ("BMC").
Press to continue or to skip:
Press to continue or to skip: S
Do you accept and agree to the terms of this Software License Agreement (Y/N)?Y
=== Control-M/Agent 9.0.00 Installation ===
=== Product Settings : Control-M/Agent parameters ===
1 : Primary Control-M/Server host:
2 : Authorized Control-M/Server hosts:
3 : Agent-to-Server port number: 7005
4 : Server-to-Agent port number: 7006
== Cancel Previous Panel / Next Panel ==
Enter command or item number you wish to change: 1
Enter a new value for Primary Control-M/Server host: controlm
=== Control-M/Agent 9.0.00 Installation ===
=== Product Settings : Control-M/Agent parameters ===
1 : Primary Control-M/Server host: controlm
2 : Authorized Control-M/Server hosts:
3 : Agent-to-Server port number: 7005
4 : Server-to-Agent port number: 7006
== Cancel Previous Panel / Next Panel ==
Enter command or item number you wish to change: 2
Enter a new value for Authorized Control-M/Server hosts: controlm
=== Control-M/Agent 9.0.00 Installation ===
=== Summary ===
You are about to install Control-M/Agent version 9.0.00
Destination Folder: /export/home/x300168
Primary Control-M/Server Host: controlm
Control-M/Server Hosts: controlm
Agent To Server Port: 7005
Server To Agent Port: 7006
Agent will be started as: Root
== Cancel Previous Panel / Install Generate an automatic installation configuration file ==
Enter command:
=== Control-M/Agent 9.0.00 Installation ===
=== Installation ===
[ Starting to unpack ]
[ Processing package: Check pre-requirements for Control-M/Agent ]
[ Processing package: Control-M/Agent files ]
\
[ Processing package: Configure Control-M/Agent ]
[ Processing package: Create softlinks for Control-M/Agent ]
[ Processing package: Start and enable Control-M Agent ]
[ Processing package: The installed-versions.txt update. ]
[ Processing package: Write uninstaller ]
=== Control-M/Agent 9.0.00 Installation ===
=== Installation Result - Success ===
Installation has completed successfully.
Note:
1. To start working with Control-M/Agent 9.0.00, you must close
the current session and open a new one.
2. If you want the agent to start automatically when the host is restarted,
implement the '/export/home/x300168/ctm/scripts/rc.agent_user' script
or replace any existing startup script with it.
For more information, refer to the Installation Guide.
Debo exportar la variableCONTROLM con el directorio donde se instalara el producto
[sun5004] NUEVO /export/home/x300168 # export CONTROLM=/export/home/x300168/ctm
Ahora si, ejecuto el script de instalacion del FixPack
[sun5004] NUEVO /export/home/x300168 # ./PAKAI.9.0.00.200_Solaris_INSTALL.BIN
Extracting files...
Succeeded in extracting files to /tmp/PAKAI.9.0.00.200_Solaris_INSTALL.BIN_root.
Initiating installer, please wait...
Starting installation of Control-M/Agent 9.0.00 Fix Pack 2
Are you sure you want to install Control-M/Agent 9.0.00 Fix Pack 2? [y/n]: y
Copy WinInstallHelper
Updating settings...
Updating configurations...
Copying files...
Create exe_900fp2 directory
Copy exe directory
Updating settings...
Moving files...
Updating settings...
Copying files...
Updating configurations...
Moving files...
Change binaries ownership
Sending ping to the server
Update Symbolic Link
Creating Job Snapshot file
Chown Job Snapshot file
Renaming Folder name...
Deleting files...
Updating configurations...
Create ctmaggetcm Symbolic Link
Change backup folder ownership
Copy WinInstallHelper To Data
Updating configurations...
Control-M/Agent 9.0.00 Fix Pack 2 installation completed successfully.
Log file:
/export/home/x300168/ctm/install/PAKAI.9.0.00.200/trace.log
[sun5004] NUEVO /export/home/x300168 #
Agregar Filesystem al Sun Cluster 3.3
La finalidad es agregar el filesystem /u27 al Cluster ( conviene hacerlo con la BD baja)
1) Descubro los discos en Solaris
cfgadm -f -c configure c(nro de controladora)
devfsadm -C ( -C significa modo Clean Up )
devfsadm ( crea special files en el /devices y logical links en /dev )
scdidadm -C ( remueve todas las referencias DID de todos los dispositivos Detachados )
scdidadm -r ( reconfiguro los DID ) (en ambos nodos )
scgdevs ( updatea el /globaldevices )
scdpm -p all:all ( chequea que este bien el did en ambos nodos del cluster, suele demorar un rato , no menos de 5 minutos dependiendo la cantidad de discos que tenga)
OJO: Los comandos arriba mencionados,Deben ser ejecutados en AMBOS nodos del cluster
vxdctl enable
agrego los discos en el Veritas ( uso por ejemplo el vxdiskadm menu opcion 1 )
# vxdisk list
# vxassist -g CCtdg make vol27 51687m layout=stripe CCtdg26 CCtdg27 CCtdg28
# cldevicegroup sync CCtdg
# newfs /dev/vx/rdsk/CCtdg/vol27
Poner el NUEVO FILESYSTEM en el /etc/vfstab de AMBOS NODOS
Conviene poner el cluster en modo mantenimiento, para poder bajar y subir la base de datos, las veces que se necesite y que no se migre el cluster.
clrg suspend cct-rg
Ahora puedo agregar el nuevo filesystem al control del cluster
clrs set -g cct-rg -p FileSystemMountPoints+="/u27" cct-hastp-rs
clrs show -v cct-hastp-rs
Luego , poner el cluster en modo normal
clrg resume cct-rg
Capturando paquetes en un puerto en particular, con el comando snoop
El equipo en cuestion, tiene una zona global y 6 zonas no globales.
Quiero saber si alguno de los containers recibe paquetes por el puerto 4175
La zona global se llama sunp008 y las no globales, sunpc18,sunpc28,sunpc38,sunpc48 y sunpc58
Desde la zona global
root@sunp008-ldm # snoop -o /usr/scripts/snoop4175.txt -vd vnet150000 port 4175
Using device vnet150000 (promiscuous mode)
0
10 ^C
Y se lee con el comando snoop –i
root@sunp008-ldm # snoop -i snoop4175.txt
1 0.00000 10.24.2.94 -> sunpc48.coneja.com.ar UDP D=4175 S=51980 LEN=422
2 378.84111 sunpc58.coneja.com.ar -> 172.18.4.174 ICMP Destination unreachable (UDP port 8122 unreachable)
3 0.00375 sunpc58.conneja.com.ar -> 172.18.4.174 ICMP Destination unreachable (UDP port 8122 unreachable)
4 50.00006 sunpc58.coneja.com.ar -> 172.18.4.174 ICMP Destination unreachable (UDP port 8122 unreachable)
5 0.00394 sunpc58.coneja.com.ar -> 172.18.4.174 ICMP Destination unreachable (UDP port 8122 unreachable)
6 49.98555 sunpc58.coneja.com.ar -> 172.18.4.174 ICMP Destination unreachable (UDP port 8122 unreachable)
7 0.00401 sunpc58.coneja.com.com.ar -> 172.18.4.174 ICMP Destination unreachable (UDP port 8122 unreachable)
8 50.37636 sunpc58.coneja.com.ar -> 172.18.4.174 ICMP Destination unreachable (UDP port 8122 unreachable)
9 0.00396 sunpc58.coneja.com.ar -> 172.18.4.174 ICMP Destination unreachable (UDP port 8122 unreachable)
10 0.70132 10.24.2.94 -> sunpc48.coneja.com.ar UDP D=4175 S=56099 LEN=279
root@sunp008-ldm #
Quiero saber si alguno de los containers recibe paquetes por el puerto 4175
La zona global se llama sunp008 y las no globales, sunpc18,sunpc28,sunpc38,sunpc48 y sunpc58
Desde la zona global
root@sunp008-ldm # snoop -o /usr/scripts/snoop4175.txt -vd vnet150000 port 4175
Using device vnet150000 (promiscuous mode)
0
10 ^C
Y se lee con el comando snoop –i
root@sunp008-ldm # snoop -i snoop4175.txt
1 0.00000 10.24.2.94 -> sunpc48.coneja.com.ar UDP D=4175 S=51980 LEN=422
2 378.84111 sunpc58.coneja.com.ar -> 172.18.4.174 ICMP Destination unreachable (UDP port 8122 unreachable)
3 0.00375 sunpc58.conneja.com.ar -> 172.18.4.174 ICMP Destination unreachable (UDP port 8122 unreachable)
4 50.00006 sunpc58.coneja.com.ar -> 172.18.4.174 ICMP Destination unreachable (UDP port 8122 unreachable)
5 0.00394 sunpc58.coneja.com.ar -> 172.18.4.174 ICMP Destination unreachable (UDP port 8122 unreachable)
6 49.98555 sunpc58.coneja.com.ar -> 172.18.4.174 ICMP Destination unreachable (UDP port 8122 unreachable)
7 0.00401 sunpc58.coneja.com.com.ar -> 172.18.4.174 ICMP Destination unreachable (UDP port 8122 unreachable)
8 50.37636 sunpc58.coneja.com.ar -> 172.18.4.174 ICMP Destination unreachable (UDP port 8122 unreachable)
9 0.00396 sunpc58.coneja.com.ar -> 172.18.4.174 ICMP Destination unreachable (UDP port 8122 unreachable)
10 0.70132 10.24.2.94 -> sunpc48.coneja.com.ar UDP D=4175 S=56099 LEN=279
root@sunp008-ldm #
martes, 23 de mayo de 2017
Ver y Crear copia de la configuracion de los LDOM's
Esta configuracion conviene mantenerla actualizada, sobre todo luego de hacer cambios sobre los Logical Domains.
root@m5-2-pdom # ldm list-spconfig
factory-default
io-domain
08072016
19052017
20052017 [next poweron]
root@m5-2-pdom #
Con el siguiente comando creo una copia de configuracion de los Logical Domains
**** Suele demorar unos 10/15 segundos
El formato que le voy a dar es dia mes año
root@m5-2-pdom # ldm add-spconfig 23052017
Chequeo para ver como quedo
root@m5-2-pdom # ldm list-spconfig
factory-default
io-domain
08072016
19052017
20052017
23052017 [current]
root@m5-2-pdom #
martes, 25 de abril de 2017
ERROR: V-3-20003: Cannot open /dev/vx/dsk... No such device or address y UX:vxfs mount: ERROR: V-3-24996: Unable to get disk layout version
Luego de hacer un vxdg import de un Diskgroup llamado VARIOS, y al intentar montar un filesystem de dicho Disk Group me salia el siguiente error :
[sun5004] / # vxdg import VARIOS
[sun5004] / # mount -F vxfs /dev/vx/dsk/VARIOS/vol04 /varios
UX:vxfs mount: ERROR: V-3-20003: Cannot open /dev/vx/dsk/VARIOS/vol04: No such device or address
UX:vxfs mount: ERROR: V-3-24996: Unable to get disk layout version
[sun5004] / #
El problema en este caso fue que los volumenes de este DiskGroup estaban en estado DISABLED.
Lo chequeo con este comando
[sun5004] / # vxprint -htg VARIOS
dg VARIOS default default 48000 1350408026.360.sun5004
dm VARIOS01 EMC1_21 auto 65536 419354368 -
dm VARIOS02 EMC1_23 auto 65536 419354368 -
dm VARIOS03 EMC1_5 auto 65536 419354368 -
dm VARIOS04 EMC1_44 auto 65536 419354368 -
dm VARIOS05 EMC1_45 auto 65536 419354368 -
dm VARIOS06 EMC1_51 auto 65536 418690064 -
dm VARIOS07 EMC1_52 auto 65536 418690064 -
v vol01 - DISABLED ACTIVE 136314880 SELECT - fsgen
pl vol01-01 vol01 DISABLED ACTIVE 136314880 CONCAT - RW
sd VARIOS03-01 vol01-01 VARIOS03 0 136314880 0 EMC1_5 ENA
v vol03 - DISABLED ACTIVE 304087040 SELECT - fsgen
pl vol03-01 vol03 DISABLED ACTIVE 304087040 CONCAT - RW
sd VARIOS04-01 vol03-01 VARIOS04 0 304087040 0 EMC1_44 ENA
v vol04 - DISABLED ACTIVE 209715200 SELECT - fsgen
pl vol04-01 vol04 DISABLED ACTIVE 209715200 CONCAT - RW
sd VARIOS05-01 vol04-01 VARIOS05 0 209715200 0 EMC1_45 ENA
La Solucion fue habilitarlos de la siguiente manera
[sun5004] / # vxrecover -g VARIOS -sbE
chequeo
[sun5004] / # vxprint -htg VARIOS
DG NAME NCONFIG NLOG MINORS GROUP-ID
dg VARIOS default default 48000 1350490026.360.psun5004
dm VARIOS01 EMC1_21 auto 65536 419354368 -
dm VARIOS02 EMC1_23 auto 65536 419354368 -
dm VARIOS03 EMC1_5 auto 65536 419354368 -
dm VARIOS04 EMC1_44 auto 65536 419354368 -
dm VARIOS05 EMC1_45 auto 65536 419354368 -
dm VARIOS06 EMC1_51 auto 65536 418690064 -
dm VARIOS07 EMC1_52 auto 65536 418690064 -
v vol01 - ENABLED ACTIVE 136314880 SELECT - fsgen
pl vol01-01 vol01 ENABLED ACTIVE 136314880 CONCAT - RW
sd VARIOS03-01 vol01-01 VARIOS03 0 136314880 0 EMC1_5 ENA
v vol03 - ENABLED ACTIVE 304087040 SELECT - fsgen
pl vol03-01 vol03 ENABLED ACTIVE 304087040 CONCAT - RW
sd VARIOS04-01 vol03-01 VARIOS04 0 304087040 0 EMC1_44 ENA
v vol04 - ENABLED ACTIVE 209715200 SELECT - fsgen
pl vol04-01 vol04 ENABLED ACTIVE 209715200 CONCAT - RW
sd VARIOS05-01 vol04-01 VARIOS05 0 209715200 0 EMC1_45 ENA
Ahora si, puedo montar los filesystems
[sun5004] / # mount -F vxfs /dev/vx/dsk/VARIOS/vol04 /varios
Esta solucion fue aportada por mi colega Hugo Albarracin ( @halbarracin )
viernes, 10 de febrero de 2017
Review de solaris para It Central Station
Aca les dejo el link de una review tecnica que escribi para la prestigiosa web IT Central Station de New York
Ver revision tecnica de Solaris
Ver revision tecnica de Solaris
lunes, 12 de diciembre de 2016
ORA-27125: unable to create shared memory segment
Al intentar crear una instancia de bd Oracle, falla con el siguiente error
ORA-27125: unable to create shared memory segment
SVR4 Error: 22: Invalid argument
El server en cuestion es un SPARC T5-2 con 8 gb de RAM, con 3 zonas, en la zona 1 estaran las bd, y en la zona 2 y zona3 la aplicacion.
Lo primero que chequeo es en la zona global, si hay errores en el /var/adm/messages, y veo el mensaje que no hay espacio en el SWAP
Dec 12
11:30:02 net1002 genunix: [ID 470503 kern.warning] WARNING: Sorry, no swap
space to grow stack for pid 29204 (discusge)
root@net1002
# zfs list
NAME
USED AVAIL REFER MOUNTPOINT
rpool
16.0G 10.3G 106K /rpool
rpool/ROOT
6.35G 10.3G 31K legacy
rpool/ROOT/s10s_u11wos_24a
6.35G 10.3G 6.35G /
rpool/dump
1.50G 10.3G 1.50G -
rpool/export
73K 10.3G 36K /export
rpool/export/home
37K 10.3G 37K /export/home
rpool/swap
8.16G 10.6G 7.91G -
root@net1002
# df -h
Filesystem size used avail capacity Mounted on
rpool/ROOT/s10s_u11wos_24a
26G 6.3G 10G
39% /
/devices
0K 0K
0K 0% /devices
ctfs
0K 0K
0K 0% /system/contract
proc
0K 0K 0K
0% /proc
mnttab
0K 0K
0K 0% /etc/mnttab
swap
1.5G 448K 1.5G
1% /etc/svc/volatile
objfs
0K 0K
0K 0% /system/object
sharefs
0K 0K
0K 0% /etc/dfs/sharetab
/platform/sun4v/lib/libc_psr/libc_psr_hwcap3.so.1
17G 6.3G 10G
39% /platform/sun4v/lib/libc_psr.so.1
/platform/sun4v/lib/sparcv9/libc_psr/libc_psr_hwcap3.so.1
17G 6.3G 10G
39% /platform/sun4v/lib/sparcv9/libc_psr.so.1
fd
0K 0K
0K 0% /dev/fd
swap
1.5G 32K 1.5G
1% /tmp
swap
1.5G 88K 1.5G
1% /var/run
rpool/export
26G
36K 10G 1% /export
rpool/export/home
26G 37K 10G
1% /export/home
rpool
26G 106K 10G
1% /rpool
/dev/md/dsk/d300
30G 27G 2.7G 91%
/export/zona3
/dev/md/dsk/d200
30G 25G 4.3G
86% /export/zona2
/dev/md/dsk/d100
30G 5.3G 24G
19% /export/zona1
Ahora voy a consultar cuanto espacio libre tiene el pool RPOOL , para saber cuanto espacio puedo agregarle al volumen SWAP ( tiene 11gb libres)
root@net1002
# zpool get all rpool
NAME
PROPERTY
VALUE
SOURCE
rpool
size 26.8G
-
rpool
capacity
58%
-
rpool
altroot
-
default
rpool
health
ONLINE
-
rpool
guid
17834260759408459067 -
rpool
version
32
default
rpool
bootfs
rpool/ROOT/s10s_u11wos_24a local
rpool
delegation
on
default
rpool
autoreplace
off
default
rpool
cachefile -
default
rpool
failmode
continue
local
rpool
listsnapshots
on
default
rpool
autoexpand
off
default
rpool
free
11.0G
-
rpool
allocated
15.8G
-
rpool
readonly off
Con este comando veo que tiene asignado 7.91gb para el volumen swap
root@net1002
# zfs get all rpool/swap
NAME
PROPERTY
VALUE
SOURCE
rpool/swap
type
volume
-
rpool/swap
creation
Thu Sep 8 13:47 2016 -
rpool/swap
used
8.16G
-
rpool/swap
available
10.6G
-
rpool/swap
referenced
7.91G
-
rpool/swap
compressratio
1.00x
-
rpool/swap
reservation
none
default
rpool/swap
volsize
7.91G
local
rpool/swap
volblocksize
1M
-
rpool/swap
checksum
off
local
rpool/swap
compression
off
local
rpool/swap
readonly
off
default
rpool/swap
shareiscsi
off
default
rpool/swap
copies
1
default
rpool/swap
refreservation
8.16G
local
rpool/swap
primarycache
metadata
local
rpool/swap
secondarycache
all
default
rpool/swap
usedbysnapshots 0
-
rpool/swap
usedbydataset
7.91G
-
rpool/swap
usedbychildren
0
-
rpool/swap
usedbyrefreservation
255M
-
rpool/swap
logbias
latency
default
rpool/swap
sync
standard
default
rpool/swap
rekeydate
Ahora voy a agrandar el tamaño del volumen swap de 8gb que tenia a 16gb
root@net1002
# zfs set volsize=16g rpool/swap
root@net1002
# zfs get all rpool/swap
NAME
PROPERTY
VALUE
SOURCE
rpool/swap
type
volume
-
rpool/swap
creation
Thu Sep 8 13:47 2016 -
rpool/swap
used
16.5G
-
rpool/swap
available
2.48G
-
rpool/swap
referenced 16.0G
-
rpool/swap
compressratio
1.00x
-
rpool/swap
reservation
none
default
rpool/swap
volsize
16G
local
rpool/swap
volblocksize
1M -
rpool/swap
checksum
off
local
rpool/swap
compression
off
local
rpool/swap
readonly
off
default
rpool/swap
shareiscsi
off default
rpool/swap
copies
1
default
rpool/swap
refreservation
16.5G
local
rpool/swap
primarycache
metadata
local
rpool/swap
secondarycache
all default
rpool/swap
usedbysnapshots
0
-
rpool/swap
usedbydataset
16.0G
-
rpool/swap
usedbychildren
0
-
rpool/swap
usedbyrefreservation
516M
-
rpool/swap
logbias
latency
default
rpool/swap
sync
standard
default
rpool/swap rekeydate
en la zona1 aumento el parametro del shared memory tambien a 16gb.( esto es opcional, si no lo seteo , y dejo en 8gb que era el tamaño que estaba anteriormente, funciona igual )
root@net1c12 # projmod -s -K "project.max-shm-memory=(priv,17179869184,deny)" user.oracle
martes, 27 de septiembre de 2016
Como Agregar placa fibra a un sparc t7-1
En este breve instructivo mostramos como insertar una placa de fibra para SAN , valido para cualquier placa pci.
La instalacion en este caso especifico corresponde a 1 Sun Storage Dual 16 Gb Fibre Channel PCIe Universal HBA, Qlogic
Para instalar o remover una placa de fibra o fiber channel, el equipo tiene que estar apagado , sin energia electrica.
Para esto vamos a conectar un cable serial a la SP del T7
Desconectamos los cables de la fuente de energia ( power cords)
Ahora comenzamos con la apertura del server, como vemos a continuacion,
Para poder insertar o remover una placa pci, es necesario abrir el mecanismo de locking que vemos a continuacion
En la vida real seria este :
Perfecto !!! Ahora a volver a armar las tapas ( top cover ), deslizamos el server hacia atras volviendo a su posicion original en el rack, presionando las trabas verdes de los laterales y conectamos los cables a las fuentes, dando energia al equipo ( power cords)
Encendemos el equipo. Nos conectamos a la SP y hacemos un start /System y un start /HOST/console
Al encenderlo, se toma varios minutos para correr un POST , con un diagnostico extendido.
miércoles, 21 de septiembre de 2016
Problem with LDMD daemon and the solution (spanish version)
In this article , we describe with my collegue Nicolas Morono, a bug with ldmd daemon and how to restore the previous configuration of the Logical Domains ( LDOMs ) using ldm-db.xml file
When we wanted assign a lun to a LDOM, we find with this trouble :
# ldm list
Failed to connect to logical domain manager: Connection refused
We check and the service ldmd is in maintenance state
# svcs -xv
svc:/ldoms/ldmd:default (Logical Domains Manager)
State: maintenance since June 2, 2016 06:36:16 PM ART
Reason: Start method exited with $SMF_EXIT_ERR_FATAL.
See: /var/svc/log/ldoms-ldmd:default.log
Impact: This service is not running.
In the /var/adm/messages it showed this errors
Jun 2 18:36:16 m5-1-pdom2 svc.startd[33]: [ID 652011 daemon.warning] svc:/ldoms/ldmd:default: Method "/opt/SUNWldm/bin/ldmd_start" failed with exit status 95.
Jun 2 18:36:16 m5-1-pdom2 svc.startd[33]: [ID 748625 daemon.error] ldoms/ldmd:default failed fatally: transitioned to maintenance (see 'svcs -xv' for details)
Jun 2 18:36:16 m5-1-pdom2 fmd: [ID 377184 daemon.error] SUNW-MSG-ID: SMF-8000-YX, TYPE: defect, VER: 1, SEVERITY: major
Jun 2 18:36:16 m5-1-pdom2 EVENT-TIME: Thu Jun 2 18:36:16 ART 2016
Jun 2 18:36:16 m5-1-pdom2 PLATFORM: SPARC-M5-32, CSN: AK00xx8x1, HOSTNAME: m5-1-pdom2
Jun 2 18:36:16 m5-1-pdom2 SOURCE: software-diagnosis, REV: 0.1
Jun 2 18:36:16 m5-1-pdom2 EVENT-ID: 889f64a0-0102-efd6-997f-8e83e7fba09a
Jun 2 18:36:16 m5-1-pdom2 DESC: A service failed - a start, stop or refresh method failed.
Jun 2 18:36:16 m5-1-pdom2 AUTO-RESPONSE: The service has been placed into the maintenance state.
Jun 2 18:36:16 m5-1-pdom2 IMPACT: svc:/ldoms/ldmd:default is unavailable.
Jun 2 18:36:16 m5-1-pdom2 REC-ACTION: Run 'svcs -xv svc:/ldoms/ldmd:default' to determine the generic reason why the service failed, the location of any logfiles, and a list of other services impacted. Please refer to the associated reference document athttp://support.oracle.com/msg/SMF-8000-YX for the latest service procedures and policies regarding this diagnosis.
Jun 2 18:40:28 m5-1-pdom2 cmlb: [ID 107833 1
We check in the svc logs
# cat /var/svc/log/ldoms-ldmd:default.log
Jun 02 18:35:16 timeout waiting for op HVctl_op_get_bulk_res_stat
Jun 02 18:35:16 fatal error: waiting for hv response timeout
[ Jun 2 18:35:16 Stopping because process dumped core. ]
[ Jun 2 18:35:16 Executing stop method (:kill). ]
[ Jun 2 18:35:16 Executing start method ("/opt/SUNWldm/bin/ldmd_start"). ]
Jun 02 18:36:16 timeout waiting for op HVctl_op_hello
Jun 02 18:36:16 fatal error: waiting for hv response timeout
[ Jun 2 18:36:16 Method "start" exited with status 95. ]
We looked at the oracle docs and came to the conclusion that there was a bug in firmware versions below 1. 14.2 which matched our environment.
We opened a service request to confirm the analyzed by us and the proposed solution was the same.
The bug is in Hypervisors lower than the version 1.14.2 .
- The short term solution is to perform a power-cycle the system.
- The solution to medium / long term is to update the system firmware to a recent version ( HypV 1.14.2 or Higher )
At this point we find that solutions involve a power cycle that involves all running LDOMS and total reboot of the machine.
We decided to perform the firmware upgrade and make the power-cycle, but we realized that the last saved settings LDOMS is old and we will lose 6 months changes in LDOMs configurations ( like creation of new LDOMs , disk assignments, allocation of network cards, etc )
The solution applied to solved this situation was as follow :
Prior to reboot the PDOM, we backup the file ldom-db.xml located in /var/opt/SUNWldm , ( this file make the Magic ) this file has all the settings that are active in PDOM regardless of whether or not you saved in the SP .
We copy this file ( ldom-db.xml ) in /usr/scripts , to use after easily without a restore from the backup
Here are the steps used
From the ilom
We make the power-cycle
stop Servers/PDomains/PDomain_2/HOST
y then
start Servers/PDomains/PDomain_2/HOST
Once we Boot the PDOM and with the LDOMs down and unbind, we take a backup of the file ldom-db.xml and disable the ldom service daemon.
root@ # ldm ls
NAME STATE FLAGS CONS VCPU MEMORY UTIL NORM UPTIME
primary active -n-cv- UART 8 16G 0.2% 0.2% 8d 2h 38m
dnet1002 active -n---- 5002 8 8G 0.5% 0.5% 5d 2h 49m
dsunt100 active -n---- 5000 48 40G 0.0% 0.0% 8d 1h 34m
dsunt200 active -n---- 5001 48 40G 0.0% 0.0% 2m
root@#
root@ # ldm stop dsunt200
LDom dsunt200 stopped
root@ # ldm unbind dsunt200
root@ # ldm stop dsunt100
LDom dsunt100 stopped
root@ # ldm unbind dsunt100
root@ # ldm stop dnet1002
LDom dnet1002 stopped
root@ # ldm unbind dnet1002
root@ # ldm ls
NAME STATE FLAGS CONS VCPU MEMORY UTIL NORM UPTIME
primary active -n-cv- UART 8 16G 0.5% 0.5% 8d 2h 40m
dnet1002 inactive ------ 8 8G
dsunt100 inactive ------ 48 40G
dsunt200 inactive ------ 48 40G
root@ #
cd /var/opt/SUNWldm
cp -p ldom-db.xml ldom-db.xml.orig
svcadm disable ldmd
##### Here we use the file stored previoulsy in /usr/scripts/, Now we overwrite the original stored in /var/opt/SUNWldm
cp -p /usr/scripts/ldom-db.xml /var/opt/SUNWldm/ldom-db.xml
# Enable the ldmd service.
svcadm enable ldmd
### We check the configuration to see if everythings is OK, bind and start of ldoms .
Then, we make an init 6 and after that .. bind and start to all ldoms like we show you next
root@ # ldm bind dsunt200
root@ # ldm start dsunt200
LDom dsunt200 started
root@ # ldm bind dsunt100
root@ # ldm start dsunt100
LDom dsunt100 started
root@ # ldm bind dnet1002
root@ # ldm start dnet1002
LDom dnet1002 started
root@ # ldm ls
NAME STATE FLAGS CONS VCPU MEMORY UTIL NORM UPTIME
primary active -n-cv- UART 8 16G 3.7% 3.7% 8d 2h 55m
dnet1002 active -n---- 5002 8 8G 0.7% 0.7% 3s
dsunt100 active -n---- 5000 48 40G 0.0% 0.0% 2s
dsunt200 active -n---- 5001 48 40G 9.1% 1.0% 2s
root@ #
PS : Please forgive my english ;-)
miércoles, 14 de septiembre de 2016
Problema con LDMD y solucion aplicada
( english version)
En este documento , describimos con mi compañero Nicolas Morono, un problema con el demonio ldmd y como lo recuperar la configuracion de los LDOMs desde el archivo ldm-db.xml
Al querer asignarle una lun a un dominio nos dio el siguiente error :
# ldm list
Failed to connect to logical domain manager: Connection
refused
chequeo y el servicio ldmd esta caido
# svcs -xv
svc:/ldoms/ldmd:default (Logical Domains Manager)
State: maintenance since June 2, 2016 06:36:16 PM ART
Reason: Start method exited with $SMF_EXIT_ERR_FATAL.
See: /var/svc/log/ldoms-ldmd:default.log
Impact: This service is not running.
en el /var/adm/messages estan estos errores registrados
Jun 2 18:36:16 m5-1-pdom2 svc.startd[33]: [ID 652011
daemon.warning] svc:/ldoms/ldmd:default: Method
"/opt/SUNWldm/bin/ldmd_start" failed with exit status 95.
Jun 2 18:36:16 m5-1-pdom2 svc.startd[33]: [ID 748625
daemon.error] ldoms/ldmd:default failed fatally: transitioned to maintenance
(see 'svcs -xv' for details)
Jun 2 18:36:16 m5-1-pdom2 fmd: [ID 377184 daemon.error]
SUNW-MSG-ID: SMF-8000-YX, TYPE: defect, VER: 1, SEVERITY: major
Jun 2 18:36:16 m5-1-pdom2 EVENT-TIME: Thu Jun 2 18:36:16 ART
2016
Jun 2 18:36:16 m5-1-pdom2 PLATFORM: SPARC-M5-32, CSN:
AK00xx8x1, HOSTNAME: m5-1-pdom2
Jun 2 18:36:16 m5-1-pdom2 SOURCE: software-diagnosis, REV:
0.1
Jun 2 18:36:16 m5-1-pdom2 EVENT-ID:
889f64a0-0102-efd6-997f-8e83e7fba09a
Jun 2 18:36:16 m5-1-pdom2 DESC: A service failed - a start,
stop or refresh method failed.
Jun 2 18:36:16 m5-1-pdom2 AUTO-RESPONSE: The service has
been placed into the maintenance state.
Jun 2 18:36:16 m5-1-pdom2 IMPACT: svc:/ldoms/ldmd:default is
unavailable.
Jun 2 18:36:16 m5-1-pdom2 REC-ACTION: Run 'svcs -xv
svc:/ldoms/ldmd:default' to determine the generic reason why the service
failed, the location of any logfiles, and a list of other services impacted.
Please refer to the associated reference document at http://support.oracle.com/msg/SMF-8000-YX
for the latest service procedures and policies regarding this diagnosis.
Jun 2 18:40:28 m5-1-pdom2 cmlb: [ID 107833 1
En el log del svc estan estos errores
# cat /var/svc/log/ldoms-ldmd:default.log
Jun 02 18:35:16 timeout waiting for op
HVctl_op_get_bulk_res_stat
Jun 02 18:35:16 fatal error: waiting for hv response timeout
[ Jun 2 18:35:16 Stopping because process dumped core. ]
[ Jun 2 18:35:16 Executing stop method (:kill). ]
[ Jun 2 18:35:16 Executing start method
("/opt/SUNWldm/bin/ldmd_start"). ]
Jun 02 18:36:16 timeout waiting for op HVctl_op_hello
Jun 02 18:36:16 fatal error: waiting for hv response timeout
[ Jun 2 18:36:16 Method "start" exited with status
95. ]
Se busco en los documentos de oracle y llegamos a la
conclusion de que habia un bug en versiones de firmware inferiores a 1.14.2 que
se correspondia con nuestro entorno.
Se abrio un service request para confirmar lo analizado por
nosotros y la solucion propuesta fue la misma.
En si, el bug es en Hypervisors inferiores a la version
1.14.2.
- La solución a corto plazo es realizar un power-cycle del
sistema.
- La solución a mediano/largo plazo es actualizar el
Firmware del sistema a una version reciente (HypV 1.14.2 o superior).
Llegado a este punto nos encontramos que las soluciones
implican un power-cycle que implica bajar todos los ldoms corriendo y el
reinicio total de la maquina.
Optamos por realizar el upgrade de firmware y al realizar el
powercycle nos damos cuenta que la ultima configuracion de ldoms guardada es
vieja y vamos a perder 6 meses de
modificaciones en las configuraciones de los ldoms.
(creaciones de ldoms nuevos, asignaciones de discos, asignacion de placas de
red, etc)
La solucion aplicada para solucionar esto fue la siguiente:
Se busco en /var/opt/SUNWldm el archivo ldom-db.xml previo a
bootear el equipo, ese archivo tiene toda lo configuracion que esta activa en
el pdom
independientemente de si esta o no guardada en la sp.
Se dejo una copia del archivo en /usr/scripts (asi no era
necesario realizar un restore del backup )
Se realiza el power-cycle desde ilom
stop
Servers/PDomains/PDomain_2/HOST
y luego
start Servers/PDomains/PDomain_2/HOST
Una vez booteado el equipo y con los ldoms bajos y unbind,
tomamos backup del archivo ldom-db.xml y deshabilitamos el demonio de ldoms.
root@ # ldm ls
NAME STATE FLAGS CONS VCPU MEMORY UTIL NORM UPTIME
primary active -n-cv- UART 8 16G 0.2% 0.2% 8d 2h 38m
dnet1002 active -n---- 5002 8 8G 0.5% 0.5% 5d 2h 49m
dsunt100 active -n---- 5000 48 40G 0.0% 0.0% 8d 1h 34m
dsunt200 active -n---- 5001 48 40G 0.0% 0.0% 2m
root@#
root@ # ldm stop dsunt200
LDom dsunt200 stopped
root@ # ldm unbind dsunt200
root@ # ldm stop dsunt100
LDom dsunt100 stopped
root@ # ldm unbind dsunt100
root@ # ldm stop dnet1002
LDom dnet1002 stopped
root@ # ldm unbind dnet1002
root@ # ldm ls
NAME STATE FLAGS CONS VCPU MEMORY UTIL NORM UPTIME
primary active -n-cv- UART 8 16G 0.5% 0.5% 8d 2h 40m
dnet1002 inactive ------ 8 8G
dsunt100 inactive ------ 48 40G
dsunt200 inactive ------ 48 40G
root@ #
cd /var/opt/SUNWldm
cp -p ldom-db.xml ldom-db.xml.orig
svcadm disable ldmd
##### aca habia backupeado y dejado en /usr/scripts/ el
archivo con la data. Ahora pisamos el original de /var/opt/SUNWldm
cp -p /usr/scripts/ldom-db.xml /var/opt/SUNWldm/ldom-db.xml
# Volvemos a habilitar el demonio.
svcadm enable ldmd
### chequeo la config , bind y start de los dominios si
quedaron ok. Aca luego de chequear la configuracion se le dio un init 6 para
que reinicie normalmente
y despues de eso se dio bind y start a los dominios.
root@ # ldm bind dsunt200
root@ # ldm start dsunt200
LDom dsunt200 started
root@ # ldm bind dsunt100
root@ # ldm start dsunt100
LDom dsunt100 started
root@ # ldm bind dnet1002
root@ # ldm start dnet1002
LDom dnet1002 started
root@ # ldm ls
NAME STATE FLAGS CONS VCPU MEMORY UTIL NORM UPTIME
primary active -n-cv- UART 8 16G 3.7% 3.7% 8d 2h 55m
dnet1002 active -n---- 5002 8 8G 0.7% 0.7% 3s
dsunt100 active -n---- 5000 48 40G 0.0% 0.0% 2s
dsunt200 active -n---- 5001 48 40G 9.1% 1.0% 2s
root@ #
Suscribirse a:
Entradas (Atom)













